Fresh daily
AI News
Latest AI tool releases, research breakthroughs, and industry news.
Older
How Amgen uses GPT-5
Learn how Amgen uses GPT-5.
Estimating worst case frontier risks of open weight LLMs
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity.
OpenAI’s new economic analysis
Analysis provides insights into ChatGPT’s impact on the economy. OpenAI also launches new research collaboration to study AI’s broader effects on the labor market and productivity.
Preparing for future AI risks in biology
Advanced AI can transform biology and medicine—but also raises biosecurity risks. We’re proactively assessing capabilities and implementing safeguards to prevent misuse.
Toward understanding and preventing misalignment generalization
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.
Disrupting malicious uses of AI: June 2025
Our latest report featuring case studies of how we’re detecting and preventing malicious uses of AI.
Introducing HealthBench
HealthBench is a new evaluation benchmark for AI in healthcare which evaluates models in realistic scenarios. Built with input from 250+ physicians, it aims to provide a shared standard for model performance and safety in health.
Expanding on what we missed with sycophancy
A deeper dive on our findings, what went wrong, and future changes we’re making.
Our updated Preparedness Framework
Sharing our updated framework for measuring and protecting against severe harm from frontier AI capabilities.
BrowseComp: a benchmark for browsing agents
BrowseComp: a benchmark for browsing agents.
New commission to provide insight as OpenAI builds the world’s best-equipped nonprofit
Already a nonprofit, and already using AI to help people solve hard problems, OpenAI aims to build the best-equipped nonprofit the world has ever seen—combining potentially historic financial resources with something even more powerful: technology that can scale human ingenuity itself.
PaperBench: Evaluating AI’s Ability to Replicate AI Research
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
Moving from intent-based bots to proactive AI agents
Moving from intent-based bots to proactive AI agents.
Early methods for studying affective use and emotional well-being on ChatGPT
An OpenAI and MIT Media Lab Research collaboration.
Detecting misbehavior in frontier reasoning models
Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.

AI tools are spotting errors in research papers
Article URL: https://www.nature.com/articles/d41586-025-00648-5 Comments URL: https://news.ycombinator.com/item?id=43295692 Points: 601 # Comments: 215
Moscow-based global news network has infected Western AI tools
Article URL: https://www.newsguardrealitycheck.com/p/a-well-funded-moscow-based-global Comments URL: https://news.ycombinator.com/item?id=43293121 Points: 167 # Comments: 107
Accelerating engineering cycles 20% with OpenAI
Accelerating engineering cycles 20% with OpenAI.
1,000 Scientist AI Jam Session
OpenAI and nine national labs bring together leading scientists for first-of-its kind event.
Deep research System Card
This report outlines the safety work carried out prior to releasing deep research including external red teaming, frontier risk evaluations according to our Preparedness Framework, and an overview of the mitigations we built in to address key risk areas.