Prediction: Meta AI scores 30/30 on Asian Physics Olympiad, signaling AI outperforming top human experts in narrow technical domains within 2-3 years
| | |

Prediction: Meta AI scores 30/30 on Asian Physics Olympiad, signaling AI outperforming top human experts in narrow technical domains within 2-3 years

The Perfect Score Nobody Wanted to Think About

Meta’s AI just scored 30 out of 30 on the Asian Physics Olympiad theoretical exam. Not a partial score. Not a near miss. A perfect score, tying with the top three human contestants in the competition.

I want you to sit with that for a moment before scrolling past it.

This is not a benchmark designed to flatter AI systems. The Asian Physics Olympiad is built to find the best physics thinkers at the high school level, anywhere on the planet. The students who sit this exam have typically trained for years, solved thousands of problems, and competed through multiple rounds of national selection. The test is designed to break most of them.

Meta’s model walked in and left with a perfect score.

Why This Is Different From Other “AI Beats Benchmark” News

We have been conditioned to shrug at these headlines. AI beats Go. AI beats protein folding. AI beats the bar exam. Each one generates a news cycle, then gets absorbed into the background hum of “AI is getting good at stuff.” I get it. The drumbeat is exhausting.

But there’s something structurally different about the Physics Olympiad result.

Purpose-built AI benchmarks have a known failure mode. They get gamed over time, intentionally or not. Training data bleeds in. Researchers optimize specifically for the target. The history of AI benchmark saturation is well documented at this point.

The Asian Physics Olympiad was not built for AI. Nobody designed it to be beatable by a language model. The problems require multi-step reasoning, physical intuition built from first principles, and the ability to construct solutions that a human expert then has to mark as correct. There is no multiple choice answer to reverse-engineer.

A general-purpose model sitting that exam and scoring 30/30 is a different category of result.

My Actual Prediction

I have been watching this space for a while and I think we are two to three years away from AI systems that routinely outperform the best human experts in specific narrow technical domains, not just on competitions, but on practical applied work.

Not all domains. Not general intelligence. I am talking about domains where the problem space is well-defined, the evaluation is unambiguous, and the solution quality is measurable. Theoretical physics problems fit that description perfectly. So does competitive programming, certain classes of mathematical proof, and narrow subfields of chemistry and materials science.

The competition result matters because it removes a comfortable excuse. People said AI could not do “real” reasoning. It just scored 30/30 on a test that demands exactly that.

What This Signals About Capability Trajectory

The pace matters here. A year ago, top models were posting respectable but imperfect scores on competition mathematics. The gap between “impressive but flawed” and “perfect score on a hard exam” closed faster than most forecasts predicted.

OpenAI’s GPT-5.6 Sol is simultaneously setting new state-of-the-art results in cybersecurity on real-world evaluation ranges, and Sam Altman noted 2.5x growth in agentic product usage in a single week. Google DeepMind is running AI agents through actual scientific hypothesis generation. These are not independent data points. They are a direction.

The trajectory is steep and it is not slowing down.

What Happens When AI Outperforms Domain Experts

This is where I think most commentary goes soft. People acknowledge the benchmark result and then pivot to “AI as a tool to augment human experts,” which is often true and sometimes the right frame. But it is not the only frame.

If a model can score 30/30 on a competition designed to identify the best human physics thinkers on Earth, the question of what “expert” means in that domain becomes genuinely uncomfortable. It does not mean human physicists stop mattering. Deep domain knowledge about which problems are worth solving, experimental intuition, and the ability to connect physics to real-world constraints will stay valuable for a long time.

But the ceiling on what AI can contribute to technical problem-solving just got revised upward, hard and fast.

The people who should be paying the most attention right now are not executives or policy writers. They are the people currently teaching and training the next generation of technical specialists. The competency map is shifting underneath them.

I do not think 2025 is the year this fully lands in public consciousness. But I am increasingly confident that 2027 looks very different from where we are standing.

🔬

Sources

#ArtificialIntelligence #MachineLearning #AIResearch #Physics #TechPredictions


Sources & Further Reading

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *