TruaceTracing the truth around AIWednesday, July 22, 2026
TRV-2026-0471Version 1 · Certified

Written 2026-07-22 03:41:27 UTC · current record

Reason for this version

Certified into the record

Canonical text (the exact bytes fingerprinted)

TRUVACE RECORD VERSION
record: TRV-2026-0471
version: 1
kind: certified
reason: Certified into the record
timestamp: 2026-07-22T03:41:27.823169Z
status: published
lens: g_space
sector: labor
headline: GDPVAL: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
dek: We introduce GDPval, a benchmark evaluating AI model capabilities on realworld economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best…
gain_title: Frontier models are approaching industry experts in deliverable quality on real-world economically valuable tasks and can perform them cheaper and faster than unaided experts when paired with human oversight.
problem_title: (none)
trace_subject: (none)
gain_reading: Frontier models are approaching industry experts in deliverable quality on real-world economically valuable tasks and can perform them cheaper and faster than unaided experts when paired with human oversight.
gain_evidence: current best frontier models are approaching industry experts in deliverable quality | to perform GDPval tasks cheaper and faster than unaided experts
problem_reading: (none)
problem_evidence: (none)
quick_read: Researchers introduced GDPval, a benchmark of real-world economically valuable tasks spanning 44 occupations and the top 9 U.S. GDP sectors, built from work of experienced industry professionals. As of January 2026, they reported frontier models improving linearly and approaching expert deliverable quality, with potential to complete tasks cheaper and faster than unaided experts when paired with human oversight.

The finding matters for labor because it moves evaluation from abstract tests to Bureau of Labor Statistics work activities that map to paid occupations. What remains uncertain is how well benchmark performance translates to unsupervised deployment, cost savings in practice, and variation across occupations not in the gold subset.
limitation: 
tag: Evidence-backed gain
key_points: GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across top 9 GDP sectors. | Tasks were constructed from representative work of industry professionals averaging 14 years of experience. | Authors open-sourced a gold subset of 220 tasks with automated grading at evals.openai.com. | Increased reasoning effort, task context, and scaffolding were found to improve model performance on the benchmark.
rundown: GDPval was built to cover the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP, using tasks drawn from industry professionals averaging 14 years of experience.

Evaluation found frontier performance improving roughly linearly over time, with current best models approaching expert deliverable quality, and showed that reasoning effort, task context, and scaffolding improve results, while authors released 220 gold tasks and a public grading service at evals.openai.com.
sources:
- peer_reviewed | SuperIntelligence - Robotics - Safety & Alignment | https://doi.org/10.70777/si.v2i4.17197 | 2026-01-21
prev: 0000000000000000000000000000000000000000000000000000000000000000
sha256
3eb281452e297846e24c7ab9de49868749113ab3e16aed5f6513a2a95dfa767a
previous
0000000000000000000000000000000000000000000000000000000000000000
Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0471 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.