Research

The closing window for public benchmarks

Twelve months out, the existing public benchmark suite will be fully saturated by every frontier model. The next generation of evaluations is going to look completely different. It is not going to be cheap.

By Telasian LabsResearch lab

The MMLU, GSM8K, HumanEval, MATH, and ARC-AGI era is ending. By mid-2027, every serious frontier model will be at or within a point of the ceiling on the existing public suites. The current generation of evaluation cannot survive the rate at which frontier capability is increasing. We have about twelve months before the public ranking becomes mostly noise.

This is not a prediction about the design of any individual benchmark. It is a structural claim about what happens when capability rises into the ceiling of every diagnostic the field has built. The benchmark stops discriminating. The press cycle stops noticing the difference between a meaningful capability gain and a marketing-grade two-point bump. The capability gap between labs widens beneath a surface that reports everyone at the same number.

This piece sketches what replaces the current generation of evaluation, what the new generation costs, and what the structural implications are for the labs that build it and the labs that depend on someone else building it.

Why this generation cannot survive

The current public benchmark suite was designed in a moment when the models being tested were below it. MMLU was hard for GPT-3. GSM8K was hard for the same generation of models. HumanEval was hard for early code models. MATH was hard for everything that came before extended reasoning was a first-class architectural element. ARC-AGI was held up as the eval that captured the gap between current systems and general intelligence. Every one of these benchmarks was designed to discriminate at a specific capability tier, and every one of them is now within reach of the frontier.

A benchmark only produces signal in the regime where the system being tested can fail it in interesting ways. Once the system clears the benchmark with margin, the failures are no longer interesting. They are noise. Two models that both score 92% on MMLU are not differentiated by that number; they are differentiated by capability dimensions the benchmark does not measure. The benchmark cannot tell you which model is better at the frontier because the benchmark was not designed to measure capability beyond its own ceiling.

The standard response has been to design harder benchmarks. MMLU-Pro followed MMLU. GPQA followed GSM8K. Humanity's Last Exam followed MATH. Each of these is, for now, harder than its predecessor. None of them are structurally different. They are all multiple-choice or short-answer evaluations against a static question bank, with public availability that makes training-distribution leakage inevitable. They will saturate on the same trajectory as the benchmarks they replaced, on a shorter timeline because the frontier is moving faster than it was when the previous generation was designed.

The benchmark designers are racing the frontier and losing. The gap between design time and saturation time is shrinking, and there is no version of this race that ends with the benchmark winning.

What replaces it

The next generation of evals does not look like the current generation. It is built around four characteristics: long horizon, domain specificity, verifiable ground truth, and adversarial design. None of these come cheap.

Long-horizon agentic tasks measure capability under extended planning, tool use, error recovery, and goal stability. They cannot be run in a few seconds of API time. They run for hours. The model is given a task that takes a competent human days to complete, a tool environment that exposes a real surface, and a budget of actions; the eval scores not on the trajectory but on the end state, evaluated against an oracle. SWE-bench is an early prototype. The next generation will be more demanding and span more domains.

Domain-specific expert evaluation requires actual experts writing rubrics that map to professional standards in their field. A practicing radiologist writes the radiology eval. A securities lawyer writes the securities-law eval. A working mathematician writes the mathematics eval at the level of an actual mathematics dissertation, not at the level of a multiple-choice exam. The rubric reflects what the field actually considers competent work, and the scoring requires either an expert reader or a separate model whose role is rubric application against the field's standards.

Verifiable ground truth means the eval has an oracle answer the rubric can check against. The oracle is sometimes a deterministic check (the test suite passes, the unit converges, the system reaches the specified state). It is sometimes an expert-written gold standard. In either case, the eval produces a signal that is not contaminated by the model's ability to produce plausible-sounding output. The bottleneck is constructing oracles, which is expensive and slow.

Adversarial design assumes the model has seen the eval's surface form during training and constructs probes that go past surface recognition. The eval is not optimized against by virtue of being trained on; it is structured so that training on it does not improve performance on the actual measured capability. This is hard, and the techniques for doing it well are themselves an active research area. Labs that take eval design seriously are publishing on adversarial design as a discipline in its own right.

The cost projection

A current-generation eval costs a few dollars to a few hundred dollars per model per pass, dominated by API time. A next-generation eval costs in the low thousands per model per pass, dominated by human expert time plus the compute to run extended agentic trajectories. The cost increase is one to two orders of magnitude.

Labs should plan budgets accordingly. If your current quarterly eval line is fifty thousand dollars across the frontier suite, the next-generation equivalent is closer to a million. The labs that do this seriously will pay it. The labs that do not will be evaluating with a broken instrument and not realizing it for years.

The cost is not just dollar cost. It is wall-clock time. A current-generation eval can run overnight. A next-generation eval, with expert scoring and long-horizon agentic trajectories, takes weeks. The release cadence at frontier labs has not yet adjusted to that reality. Releases continue to be timed against the schedule of the current-generation eval, which is going to produce systematically optimistic readings as the capability that matters most moves out of that eval's range.

Why expert time is the binding constraint

Compute scales. Synthetic data partially scales. Expert time does not scale. A senior practitioner in a domain can write a few well-constructed eval items per week if the work is being done carefully. Scaling an expert eval to the breadth the frontier requires means recruiting from a population that is small to begin with, training them on what good eval design looks like, and paying them at rates that compete with their existing professional work. None of these are quick. The labs that build serious eval infrastructure are effectively building a small research organization adjacent to their model organization, and that adjacency is itself a competitive moat.

The strategic implication

Closed labs with deeper pockets will widen the evaluation moat. The discriminating evals will be private. The public ranking will keep saying everyone is at the ceiling while the actual capability gap between frontier labs grows out of sight. Open-source models will lag specifically in the dimensions that matter most because the rigorous evaluation infrastructure to surface their gaps will be locked behind closed doors.

This is the next form of the open versus closed asymmetry, and it will be more durable than the current weight access debate. Weights you can release. Evaluation infrastructure you can keep proprietary indefinitely, and the rational thing to do is exactly that. Plan for the information asymmetry to grow.

The harder consequence is for everyone who depends on someone else's evaluation to make decisions about model capability. Procurement teams choosing between frontier offerings will be navigating with a broken compass. Regulators trying to set release criteria for high-capability models will be relying on the same broken compass. Independent researchers trying to study frontier capability will not have access to the evals that would actually surface what they are studying. The downstream effects of evaluation infrastructure consolidating inside frontier labs are not yet being seriously discussed in public, and they are larger than the weight-access discussion that has absorbed most of the attention.

What to do now

If you are a frontier lab, the budget for evaluation needs to grow by an order of magnitude over the next twelve months, and the infrastructure to support it needs to be built before it is needed rather than after the public ranking stops being useful. The labs that get caught flat-footed will spend the second half of 2027 building this in a hurry while shipping models on systematically optimistic readings.

If you are a downstream user of frontier capability, build your own eval. The construction is not exotic. Pick three tasks that matter to your work. Write your own ground truth for them. Run the eval quarterly against whatever the current frontier offers. Your eval will not be as rigorous as a frontier-lab internal eval, but it will tell you something the public benchmark stopped telling you a year ago: which model is actually best for your work.

If you are a policy actor trying to design release criteria, the current technical capacity to design evaluations at the level the frontier requires does not yet exist outside frontier labs themselves. Building that capacity, in a body that can credibly hold private evals and apply them to release decisions, is a second-order intervention worth far more attention than it currently receives. The window in which that capacity could be built ahead of the eval gap widening is open now and closing fast. After it closes, the policy conversation is structurally downstream of decisions the labs make in private, and the policy actors will be operating on the same broken compass as everyone else.

About the author

Telasian Labs

Research lab

Telasian Labs is a frontier-AI research lab. The lab studies the foundational layer of modern artificial intelligence: how frontier models are built, how LLMs and agents are designed, trained, and governed. Published work analyzes frontier developments and the science and infrastructure of intelligence itself.