GDPVAL: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
We introduce GDPval, a benchmark evaluating AI model capabilities on realworld economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best…
Frontier models are approaching industry experts in deliverable quality on real-world economically valuable tasks and can perform them cheaper and faster than unaided experts when paired with human oversight.
Evidence
- Peer-reviewedSuperIntelligence - Robotics - Safety & Alignment2026-01-21
How should this claim be treated?
Truvace Impact Record TRV-2026-0471, v1: “GDPVAL: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.” Truvace, 2026-07-22. /record/TRV-2026-0471 (accessed at citation time). sha256 3eb281452e297846…
Calibration history
Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.
Certified into the record
How to verify without trusting this page
Fetch the canonical text of any version from /api/record/TRV-2026-0471 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.
ace