Fresh daily
AI News
Latest AI tool releases, research breakthroughs, and industry news.
Older
Introducing the SWE-Lancer benchmark
Can frontier LLMs earn $1 million from real-world freelance software engineering?
Using OpenAI o1 for financial analysis
Rogo scales AI-driven financial research with OpenAI o1
Understanding complex trends with deep research
How OpenAI deep research helps Bain & Company understand complex industry trends.
Strengthening America’s AI leadership with the U.S. National Laboratories
OpenAI’s latest line of reasoning models will be used by nation’s leading scientists to drive scientific breakthroughs.
Operator System Card
Drawing from OpenAI’s established safety frameworks, this document highlights our multi-layered approach, including model and product mitigations we’ve implemented to protect against prompt engineering and jailbreaks, protect privacy and security, as well as details our external red teaming efforts, safety evaluations, and ongoing work to further refine these safeguards.
Trading inference-time compute for adversarial robustness
Trading Inference-Time Compute for Adversarial Robustness
Deliberative alignment: reasoning enables safer language models
Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.
OpenAI o1 System Card
This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teaming and frontier risk evaluations according to our Preparedness Framework.
Morgan Stanley is shaping the future of financial services
Morgan Stanley uses AI evals to shape the future of financial services
Advancing red teaming with people and AI
Advancing red teaming with people and AI
Data-driven beauty and creativity with ChatGPT
Data-driven beauty: How The Estée Lauder Companies unlocks insights with ChatGPT
Introducing SimpleQA
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Simplifying, stabilizing, and scaling continuous-time consistency models
We’ve simplified, stabilized, and scaled continuous-time consistency models, achieving comparable sample quality to leading diffusion models, while using only two sampling steps.
OpenAI and the Lenfest Institute AI Collaborative and Fellowship program
OpenAI and the Lenfest Institute AI Collaborative and Fellowship program
Evaluating fairness in ChatGPT
We've analyzed how ChatGPT responds to users based on their name, using AI research assistants to protect privacy.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
Creating agent and human collaboration with GPT 4o
Altera uses GPT-4o to build a new area of human collaboration
Using GPT-4 to improve teaching and learning in Brazil
Improving teaching and learning in Brazil
Learning to reason with LLMs
Answering quantum physics questions with OpenAI o1
Quantum physicist Mario Krenn uses OpenAI o1 to help answer life's biggest questions.