Demystifying evals for AI agents
Demystifying evals for AI agents...
Code execution with MCP: Building more efficient agents...
How To Evaluate Voice Agents Execution Outcomes And Experience...
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning an...
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence...
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and cha...
Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, ...
How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, netwo...
Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clea...
OpenAI introduces three Academy courses that help people build practical AI skills, create repeatable workflows, and app...
OpenAI plans to acquire Ona to expand Codex with secure, persistent cloud environments, enabling long-running AI agents ...
Learn how Endava is using AI agents, ChatGPT Enterprise, and Codex to accelerate software delivery, automate workflows, ...
OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new path to build with OpenAI ...
See how OpenAI, Thrive, and Crete built a self-improving tax agent with Codex, automating filings, improving accuracy, a...
Warp uses GPT-5.5 and OpenAI models to coordinate coding agents across local, cloud, and open-source development workflo...
OpenAI and Dell partner to bring Codex to hybrid and on-premise environments, helping enterprises deploy AI coding agent...
Databricks uses GPT-5.5 for enterprise agent workflows after the model set a new state of the art on the OfficeQA Pro be...
Preview a new personal finance experience in ChatGPT for Pro users in the U.S. Securely connect your financial accounts ...
Parloa leverages OpenAI models to power scalable, voice-driven AI customer service agents, enabling enterprises to desig...
OpenAI’s B2B Signals research shows how frontier enterprises deepen AI adoption, scale Codex-powered agentic workflows, ...
OpenAI and PwC are partnering to help enterprises use AI agents to automate finance workflows, improve forecasting, stre...
Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents...
Learn how Symphony, an open-source spec for Codex orchestration, turns issue trackers into always-on agent systems—boost...
Workspace agents in ChatGPT are Codex-powered agents that automate complex workflows, run in the cloud, and help teams s...
Learn how to build and use workspace agents in ChatGPT to automate repeatable workflows, connect tools, and streamline t...
The updated Codex app for macOS and Windows adds computer use, in-app browsing, image generation, memory, and plugins to...
OpenAI introduces GPT-Rosalind, a frontier reasoning model built to accelerate drug discovery, genomics analysis, protei...
OpenAI updates the Agents SDK with native sandbox execution and a model-native harness, helping developers build secure,...
Cloudflare brings OpenAI’s GPT-5.4 and Codex to Agent Cloud, enabling enterprises to build, deploy, and scale AI agents ...
CyberAgent uses ChatGPT Enterprise and Codex to securely scale AI adoption, improve quality, and accelerate decisions ac...
Gradient Labs uses GPT-4.1 and GPT-5.4 mini and nano to power AI agents that automate banking support workflows with low...
A New Framework for Evaluating Voice Agents (EVA)...
GPT-5.4 mini and nano are smaller, faster versions of GPT-5.4 optimized for coding, tool use, multimodal reasoning, and ...
How ChatGPT defends against prompt injection and social engineering by constraining risky actions and protecting sensiti...
Codex Security is an AI application security agent that analyzes project context to detect, validate, and patch complex ...
By combining rigorous model evaluation, full-platform use of OpenAI, and agent workflows, Balyasny is reinventing invest...
Introducing GPT-5.4, OpenAI’s most most capable and efficient frontier model for professional work, with state-of-the-ar...
Stateful Runtime for Agents in Amazon Bedrock brings persistent orchestration, memory, and secure execution to multi-ste...
OpenAI and Pacific Northwest National Laboratory introduce DraftNEPABench, a new benchmark evaluating how AI coding agen...
OpenAI and Paradigm introduce EVMbench, a benchmark evaluating AI agents’ ability to detect, patch, and exploit high-sev...
Introducing GPT-5.3-Codex-Spark—our first real-time coding model. 15x faster generation, 128k context, now in research p...
By Ryan Lopopolo, Member of the Technical Staff...
GPT-5.3-Codex is a Codex-native agent that pairs frontier coding performance with general reasoning to support long-hori...
GPT‑5.3-Codex is the most capable agentic coding model to date, combining the frontier coding performance of GPT‑5.2-Cod...
Introducing the Codex app for macOS—a command center for AI coding and software development with multiple agents, parall...
On February 13, 2026, alongside the previously announced retirement of GPT‑5 (Instant, Thinking, and Pro), we will reti...
A technical deep dive into the Codex agent loop, explaining how Codex CLI orchestrates models, tools, prompts, and perfo...
Discover how Higgsfield gives creators cinematic, social-first video output from simple inputs using OpenAI GPT-4.1, GPT...
ServiceNow expands access to OpenAI frontier models to power AI-driven enterprise workflows, summarization, search, and ...
How Netomi scales enterprise AI agents using GPT-4.1 and GPT-5.2—combining concurrency, governance, and multi-step reaso...
Tolan built a voice-first AI companion with GPT-5.1, combining low-latency responses, real-time context reconstruction, ...
OpenAI is strengthening ChatGPT Atlas against prompt injection attacks using automated red teaming trained with reinforc...
OpenAI introduces a new framework and evaluation suite for chain-of-thought monitorability, covering 13 evaluations acro...
OpenAI shipped Sora for Android in 28 days using Codex. AI-assisted planning, translation, and parallel coding workflows...
GPT-5.2 is our most advanced frontier model for everyday professional work, with state-of-the-art reasoning, long-contex...
Introducing GPT-5.1-Codex-Max, a faster, more intelligent agentic coding model for Codex. The model is designed for long...
This system card outlines the comprehensive safety measures implemented for GPT‑5.1-CodexMax. It details both model-leve...
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss mode...
Discover how SafetyKit leverages OpenAI GPT-5 to enhance content moderation, enforce compliance, and outpace legacy safe...
Built with OpenAI o3, o3-Pro, GPT-4.1, and GPT-5, Basis’ AI agents help accounting firms save up to 30% of their time an...
We’re releasing gpt-oss-120b and gpt-oss-20b—two state-of-the-art open-weight language models that deliver strong real-w...
Discover how Outtake uses GPT-4.1 and OpenAI o3 to power AI agents that detect and resolve digital threats 100x faster t...
As part of our Executive Function series, Model ML CEO Chaz Englander discusses how AI-native infrastructure and autonom...
Retell AI is transforming the call center with AI voice automation powered by GPT-4o and GPT-4.1. Its no-code platform e...
Unify, an AI-powered GTM platform, uses OpenAI’s o3, GPT-4.1, and CUA to automate prospecting, research, and outreach. W...
We are replacing the existing GPT-4o-based model for Operator with a version based on OpenAI o3. The API version will re...
CodeRabbit uses OpenAI models to revolutionize code reviews—boosting accuracy, accelerating PR merges, and helping devel...
Codex is a cloud-based coding agent. Codex is powered by codex-1, a version of OpenAI o3 optimized for software engineer...
Watch hands-on demos of the lastest in ChatGPT for Business: o3, image generation, enhanced memory, and internal knowled...
OpenAI o3 and OpenAI o4-mini combine state-of-the-art reasoning with full tool capabilities—web browsing, Python, image ...
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research....
This report outlines the safety work carried out for the OpenAI o3-mini model, including safety evaluations, external re...
This report outlines the safety work carried out prior to releasing OpenAI o1 and o1-mini, including external red teamin...
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering....
Color Health is working with OpenAI to pioneer a new way of accelerating cancer patients’ access to treatment. Their new...
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion mo...
We’re developing a blueprint for evaluating the risk that a large language model (LLM) could aid someone in creating a b...
We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal mo...
We trained a neural network to play Minecraft by Video PreTraining (VPT) on a massive unlabeled video dataset of human M...