GDPVal - Measuring the Performance of AI on Real World Tasks
AI Companies are racing to make AI models better at doing "economically valuable" work. OpenAI recently launched a new benchmark specifically designed to evaluate this called GDPVal. As they say:
"Our mission is to ensure that artificial general intelligence benefits all of humanity. As part of our mission, we want to transparently communicate progress on how AI models can help people in the real world. That’s why we’re introducing GDPval: a new evaluation designed to help us track how well our models and others perform on economically valuable, real-world tasks. We call this evaluation GDPval because we started with the concept of Gross Domestic Product (GDP) as a key economic indicator and drew tasks from the key occupations in the industries that contribute most to GDP."
We commend OpenAI for publishing benchmarks despite the fact that their own model is not at the top of the leaderboard. In the data shared here, Claude Opus 4.1 was the best at GDPVal. Please read the full post.
Our questions for your consideration:
Is optimizing for GDP the same as optimizing for human flourishing?
They say AI helps with "repetitive, well-specified tasks." But isn't learning through repetition how humans develop judgment? What cognitive muscles atrophy when we outsource the reps?
If AI gets better by training on expert work, and experts get replaced by AI, who creates the training data for the next generation? Are we eating our own tail?
OpenAI says they want everyone on the "up elevator" - but elevators only go up if there's somewhere to go. What floor are we actually heading to?
If Opus 4.1 is already so good, why have we decided to employ HE-2 and HE-3?
