GDPVAL: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
We introduce GDPval, a benchmark evaluating AI model capabilities on realworld economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best…
The maternal physician : a treatise on the nurture and management of infants, from the birth until two years old : being the result of sixteen years' experience in the nursery : illustrated by extracts from the most approved medical authors by American matron. Public domain
Researchers introduced GDPval, a benchmark of real-world economically valuable tasks spanning 44 occupations and the top 9 U.S. GDP sectors, built from work of experienced industry professionals. As of January 2026, they reported frontier models improving linearly and approaching expert deliverable quality, with potential to complete tasks cheaper and faster than unaided experts when paired with human oversight.
The finding matters for labor because it moves evaluation from abstract tests to Bureau of Labor Statistics work activities that map to paid occupations. What remains uncertain is how well benchmark performance translates to unsupervised deployment, cost savings in practice, and variation across occupations not in the gold subset.
- GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across top 9 GDP sectors.
- Tasks were constructed from representative work of industry professionals averaging 14 years of experience.
- Authors open-sourced a gold subset of 220 tasks with automated grading at evals.openai.com.
- Increased reasoning effort, task context, and scaffolding were found to improve model performance on the benchmark.
Frontier models are approaching industry experts in deliverable quality on real-world economically valuable tasks and can perform them cheaper and faster than unaided experts when paired with human oversight.
The rundown
GDPval was built to cover the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP, using tasks drawn from industry professionals averaging 14 years of experience.
Evaluation found frontier performance improving roughly linearly over time, with current best models approaching expert deliverable quality, and showed that reasoning effort, task context, and scaffolding improve results, while authors released 220 gold tasks and a public grading service at evals.openai.com.
Sources
- Peer-reviewedSuperIntelligence - Robotics - Safety & Alignment2026-01-21
How should this claim be treated?
ace
The debate