TechFlow Logo
Login/ Sign up
ETH Gas
Gwei
Fear
gas
A Brief History of AI Victories: Wherever There Is a Rating System, There Is AI Invasion

A Brief History of AI Victories: Wherever There Is a Rating System, There Is AI Invasion

2026.07.23
Share

TechFlow Selected TechFlow Selected

techFlow

A Brief History of AI Victories: Wherever There Is a Rating System, There Is AI Invasion

When a field establishes scoring criteria, it sets a countdown for its own conquest.

2026.07.23 - 03:46:52
AI
When a field establishes scoring criteria, it sets a countdown for its own conquest.

Author: Robonaissance

Compiled by: TechFlow

TechFlow Editor's Note: Go, chess, protein folding—why are these fields, which seem to require human intuition and creativity, being conquered by AI one after another? This article reveals an overlooked pattern: What determines whether AI can surpass humans is not task difficulty, but whether success can be scored. When a field establishes a scoring standard, it sets a countdown for its own conquest.

Selected from the "Whatever You Ask For" series, about the one job machines cannot do for us.

One in Ten Thousand

On March 10, 2016, in a hotel in Seoul, a machine made a move that no one in the room could understand, while its opponent was not present at the time.

Lee Sedol went out to smoke. He was 33 years old at the time, holding 18 world championship titles, recognized as the strongest Go player of that generation. He lost the first game the day before, and entering the second game, his pre-match confidence had tempered significantly, and his mindset was more cautious. While he was outside, AlphaGo chose move 37, and DeepMind researcher Aja Huang stood in for the machine, quietly placing the stone on the board.

The stone landed on the fifth line.

To those who don't understand Go, this meant nothing. To those who do, it was akin to nonsense. In the opening and midgame, positions that far from the edge are considered inefficient, gaining neither territory nor defending ground, equivalent to giving points to the opponent for free. This is the kind of move a teacher would correct in a beginner. In the commentary booth, top professional player Michael Redmond, while broadcasting the match, initially thought there was an error with the board signal. He picked up a stone, then put it down.

Lee Sedol returned and sat down, staring at the board for a long time without moving. Accounts vary from about 12 minutes to 15 minutes. Fan Hui, the European champion defeated by AlphaGo five months prior, was watching the match in the building. His judgment after staring for a long time was: "This is not a human move. I've never seen anyone play like this."

He was right, and right in a quantifiable way. Inside DeepMind's system there was an estimation mechanism that, by learning from a large number of human games, estimated the probability of a human making a certain move in any given position. For move 37, this estimated value was about one in ten thousand.

The machine played it anyway, and this move won the entire game.

There is a comfortable interpretation of this story, saying that the machine studied human masters very hard and finally caught up with the strongest among them. This interpretation is wrong; the number one in ten thousand explains exactly why. A system that learns to imitate human moves has a ceiling, and that ceiling is human moves. What happened in Seoul was that a system was no longer constrained by that ceiling, because the goal it was given to pursue was not the approval of human masters, but winning the game.

This distinction points in a direction more useful than admiration. If you want to know which human activities have already lost in terms of pure capability, and which are next, the question to ask is not how difficult the activity is, how much creativity it requires, or how much intuition. The question is far more boring than that.

The question is, can success be scored.

Game After Game

Chess lost 19 years ago, and it lost in a completely different way.

Deep Blue defeated Kasparov in 1997, using an architecture that was essentially a massive act of human expression. Its evaluation function—the component that looks at a position and gives a number for how good it is—was assembled and tuned with the help of grandmaster consultants. Human chess knowledge was painstakingly extracted from human players and written into the machine. What the machine added was search capability: looking ahead more moves faster, tireless, not losing pieces due to lapses in attention.

This was a real victory and should count as one. But note its form. Deep Blue's understanding of chess was human understanding. Its advantage was speed. If you ask it why it evaluated a position that way, the answer ultimately falls on someone who told it to do so.

20 years later, DeepMind released a system called AlphaZero, and the form changed.

AlphaZero was given the rules of chess. It was not given an opening book—the directory of studied first moves that every serious program relies on. It was not given endgame tablebases. It was not given a single human game. It played against itself, starting from random moves, adjusting itself based on what wins.

After about four hours, DeepMind estimated its rating exceeded the strongest traditional program at the time, Stockfish 8. After about nine hours, it played one hundred games against Stockfish under time controls, winning 28, losing none, and drawing the remaining 72. The paper reported that within 24 hours, it reached superhuman levels in chess, shogi, and Go.

These numbers are often cited without the other half of the narrative, and the other half is important. The self-play games were generated on five thousand first-generation Tensor Processing Units, with another sixty-four second-generation units training the network, all running in parallel. Four hours of wall-clock time was four hours of compute—amounts that no individual, or even a few institutions, could assemble. The achievement was not that learning chess is easy. The achievement was that when you have that much compute, human knowledge is no longer what you need.

Kasparov had more reason than anyone alive to take this result as directed at himself, yet he wrote a commentary in "Science", very generously. His observation was that AlphaZero did not play the kind of dry, cautious, draw-biased chess that everyone thought a perfect machine would play. It preferred activity over material, sacrificing pieces for positions that looked risky in his eyes. He pointed out that traditional programs carry the priorities and biases of the people who wrote them. AlphaZero wrote itself. He concluded that its style therefore reflected something more real than programmer taste, and observed that it checked far fewer positions per second than the world's best traditional engines, yet won.

This last detail is worth remembering. Traditional engines search many more possibilities per second, yet lost. AlphaZero's advantage was not looking harder, but knowing where to look, and this knowledge was assembled from millions of self-play games and a rule about who wins, nothing else.

The shogi results were even stranger to those qualified to read them. Shogi is a Japanese board game where captured pieces return to the hand of the captor, making position volatility exceed chess, and it also has centuries of accumulated theory on how to protect the king's safety. Strong players who watched the machine's games reported that the opening violated known theory, with the king running to the center of the board at moments when every book said it should retreat to the corner. These games were almost unintelligible to trained eyes, but they won.

Putting the two systems side by side, the pattern is unsettlingly clear. Deep Blue was human knowledge plus machine speed. AlphaZero was machine speed plus the definition of victory, with human knowledge deliberately removed. The version with human knowledge removed was better.

All this does not mean the 2016 machine was perfect, and the same Seoul match contained evidence.

In the fourth game, trailing three games and fighting for dignity, Lee Sedol wedged a stone between AlphaGo's two groups in the center of the board. It was later called the Divine Move. AlphaGo's own estimate of the probability that a human would play this move was also around one in ten thousand, strangely symmetric. The system failed to cope. Its evaluation of its own winning probability collapsed over the next few moves, its play deteriorated, and it lost the game.

So in March 2016, there was still a hole in the machine, one that a human found under the gaze of the whole world and under maximum pressure. This is worth stating explicitly, rather than quietly glossing over. Also worth saying is what happened later: such holes were steadily patched, successors to that system no longer lost such games, and top engines haven't lost a single game to humans in serious occasions for a long time. The 2016 result was a snapshot of a transition, not a permanent balance of power.

From these boards to everything else, what transferred was not victory itself, but the mechanism. Chess and Go were destined to fall first, for an embarrassingly simple reason. In games, success is defined by rules in a completely precise way. You win or you don't, scoring is free, instant, and indisputable. A system can play forty million games against itself in a weekend, obtaining forty million unambiguous verdicts.

This suggests where to look next. Not those simple tasks, but tasks that come with their own scoreboards.

The Fifty-Year Problem

Proteins are chains of amino acids that fold into complex three-dimensional shapes the instant they are manufactured. Shape determines what a protein does. Sequence determines shape. Deriving the latter from the former has been an open problem in biology since the early 1970s, and open in a way that resisted all attacks: from first-principles physical simulations, statistical analysis of evolutionary kinship, to structural intuition accumulated over decades, nothing worked.

In 1994, a group of researchers led by John Moult did something about this fact: everyone in the field was claiming progress, but no one could verify it.

They created an exam. Since then, every two years, Critical Assessment of Structure Prediction (CASP) takes protein structures just determined by experiment, in some cases still being determined, and publishes the sequences to the world. Any team can submit predictions. No one can access the answers, because for some targets the answers do not exist yet. When the experimental structures come out, predictions are scored against them.

Scoring uses a metric called Global Distance Test, from 0 to 100, roughly understandable as the percentage of the chain's final positions close enough to actual positions. Moult has said that about 90 points is informally considered equivalent to laboratory-determined structures. For most of the assessment's history, the best predictions hovered around 60 points.

At CASP14 in 2020, a participant registered as Group 427 achieved a median score of 92.4 across all targets. In the hardest category—targets with no useful structural kinship to rely on—the median score was 87.0. The average error was about 1.6 Angstroms, roughly equivalent to the width of an atom.

Moult has served as chair since the competition began, and he took the audience through the history of the competition before showing the chart. The chart told the whole story in one image. Lines crawled around 60 points for twenty years, then one line stood where no other line had ever been.

Group 427 was AlphaFold2. Moult announced that for single protein chains, the problem was solved.

The qualifier should be there, in every honest narrative. Single chains are not all of protein science. How proteins assemble into complexes, how they move, how they behave inside living cells, these are still open. The closed grand challenge was a specific, precisely stated challenge.

And this precision is exactly the point of telling this story here. What made structure prediction fall was not that it was originally simple, because half a century of failure indicated the opposite. What made it fall was that in 1994 the field built itself a scoreboard.

Look at what CASP provides as engineering rather than science. It provides an unambiguous measure of success, applicable to any prediction, computable within seconds. It provides a benchmark fact that cannot be cheated, because the answers are physically determined by others in the lab. It provides a stream of new problems generated continuously on a fixed schedule. It provides thirty years of historical attempts that have been scored, that is, a graded record of what counts as better and what counts as worse in this field.

A field that did all this inadvertently prepared its problems for automated attack. The scoreboard is the prerequisite. Everything else is engineering and compute, and both have been getting cheaper every year for a long time.

This generalizes with unsettling ease. Wherever a discipline reaches benchmark consensus, a target is posted there. Machine translation has scoring benchmarks. Speech recognition has scoring benchmarks. Image classification has a collection of one million labeled photos and an annual competition, and everyone who has seen that competition knows how the story ends. The pattern is not that hard things fall first, or simple things fall first. The pattern is that measured things fall first.

This raises an obvious question for everything not yet measured.

Scoring the Unscoreable

The last line of defense should have been those things that cannot be scored.

Writing is the standard example. No program can take a paragraph and return a number saying how good it is. Two competent editors disagree. The same editor disagrees with themselves on different days. The quality of language is entangled with context, audience, purpose, and taste, none of which can be reduced to a measurement. If machines need a scoreboard, and language has no scoreboard, then language is safe.

This argument holds. What happened to it is the most important reason in this entire narrative.

The field did not find an objective measure of good writing. There is still no such thing. What it did was manufacture a score using the only available material, which is human judgment itself.

The history of this method is older than most people think. In 2008, Knox and Stone described a system named TAMER, where humans observe an agent's behavior and give evaluations, which are used to train a model to predict what humans would say. Then the model rather than humans provides the training signal. In 2017, Christiano and colleagues published work regarded by most as the direct origin of current practice, applying this idea to agents in Atari games. In 2019, Ziegler and colleagues applied it to language models, and most terminology used today was established in that paper. In 2022, Ouyang and colleagues demonstrated this method at industrial scale on instruction-following tasks, and shortly thereafter, most people on Earth encountered these results without knowing it.

Its core mechanism is worth understanding in detail, because it is much simpler than its reputation sounds.

A person sits in front of two texts, both generated by the model for the same prompt. The task is not to score them. Scoring is exactly what humans do poorly, because one person's seven is another person's five, and the same person is not stable within an afternoon. The task is just to say which one is better. This is a judgment humans can make reliably, and the mathematical method to turn a bunch of such comparisons into a consistent scale is older than the fields that use it, dating back to Bradley and Terry's work on paired comparisons in 1952.

Consider the working conditions for making such judgments, because they ultimately have an impact. This person must make many such judgments per hour, for pay, against a written guide explaining what the client considers a better answer. Some pairs are almost identical. Some involve topics the person knows nothing about, in which case the more confident, better-organized of the two answers is often selected, regardless of whether it is more accurate. This is not a criticism of these people. This is a description of what anyone would do under such conditions, and the resulting choices are the raw material.

Collect enough of these choices, and you can train a second model, whose task is to predict which of the two texts humans would prefer. This second model outputs a number. And a number is the only thing that has been missing all along.

Look closely at what is established here, because it is easy to overlook. The score being optimized is not a fact about the world. It is about our model. It is a compressed, learned imitation of the judgments of a specific group of people, hired at a specific time, working under specific instructions, feeling tired in the afternoon like everyone else. CASP scores predictions based on physical reality determined in the lab, and this system scores text based on a statistical portrait of human approval.

This portrait is useful. But it is also necessarily an approximation, there are differences between it and the object depicted, and no one has fully figured out these differences. Anything about our actual preferences that our preference model fails to capture is not in the objective, therefore will not be optimized, therefore is subject to any circumstance—and powerful optimizers have no reason to protect this quantity.

This is a fact about construction, stated here without making any claims about how bad the consequences are. The point at hand is narrower and harder to refute. The category of things that cannot be scored is smaller than it appears. If a field resists measurement, a measurement can be constructed from human preferences and moved forward. This line of defense is only effective when no one thinks to manufacture a score.

What Falls Next

This leaves a practical question, and a usable answer.

Choose any activity you like: a job, a craft, a profession, a task you spent years learning to do well. Ask three questions.

First, can success and failure be reliably distinguished? Not perfectly, not by everyone, but consistently enough for competent judges to agree most of the time. Games pass this easily. Protein structure prediction passes this because the lab will eventually produce an answer. Writing does not pass in an absolute sense, but passes in a relative sense, because people can usually say which of two attempts is better.

Second, can such judgments be produced at low cost and scale? This question determines timing rather than possibility. A free and instant judgment, like in games, means the system can generate millions of training data itself. A judgment requiring a lab is slower but still feasible, with thirty years of scoring attempt archives. A judgment requiring paid humans to read text is expensive, which is exactly why so much effort has been invested in training models to imitate these humans and remove them from the loop.

Third, this determines not when a field falls, but what happens after: is this score really what you want? The score in games is what you want, because in games, winning by definition is the whole point. Protein structures measured by experiment are very close to what you want. A learned model of annotator preferences is not what you want. It is a portrait of what you want, and there is a gap between portrait and face.

Test your own work with these three questions. Most people find the first answer comes quickly and is unsettling. Most of what we call skills, at least in a relative sense, can be scored when competent people look at two attempts side by side. The second question is where the real uncertainty lies, because cost and scale are moving targets, and they have been moving in one direction for a long time. The third question is almost no one asks, and it determines what your field looks like on the other side.

It is worth doing this exercise slowly on some ordinary things. Take radiology. Can success and failure be distinguished? Yes, and precisely, because diagnoses will eventually be confirmed or refuted by the patient's condition. Can judgments be produced at low cost and scale? Basically yes, because hospitals have been accumulating scoring samples in the form of images and confirmed results for decades. Is this score what you want? Here the answer becomes complex, because what is scored is consistency with recorded diagnosis, while what is wanted is the patient doing well, these two things agree most of the time, but diverge precisely in the most important cases.

Now take something that looks safer. Take management. The first question is already hard, because competent people disagree on whether a given manager is excellent, and disagreements do not resolve quickly. The second is harder, because the results of management decisions arrive years later, entangled with everything else that happens. The third is hardest, because any alternative metric for good management proposed by anyone, from retention rates to engagement scores to output per capita, is obviously not the thing itself. According to this diagnosis, management is not safe because it is profound. It is unmeasured, which is a different, less flattering protection.

Machines have already occupied the optimization axis. On any clearly specified objective, given enough compute, finding good moves is no longer a contest, and the results are often not only stronger than ours, but stranger, reaching places our traditions teach us not to look. This is not prediction. This is a description of chess, Go, shogi, protein structure, and more and more things that were once on the list of only humans can do.

What remains is the objective itself. Someone decides what the score is. In each of the above cases, that decision took an afternoon, and every extraordinary result stems from it, is faithful to it, and is indifferent to anything it omits.

So the last question is worth pondering. In your work, is there already a score? If so, who wrote it, and did they think of you when they wrote it?

This is the second part of the "Whatever You Ask For" series, about what machines ultimately need us for, and how badly we are doing.

Join TechFlow official community to stay tuned

Add to Favorites
Share to Social Media

Related Articles

2026.07.23

AI is transitioning from a "tool" that helps you work to a "labor market" that generates income for you.

Every major technological revolution gives rise to a new generation of entrepreneurs.

AI is transitioning from a "tool" that helps you work to a "labor market" that generates income for you.
2026.07.23

Goldman Sachs Research Report Analysis: Momentum Unwinding Shocks Global Stock Markets, AI Spending Boom Conceals Hidden Risks

Before the efficiency gains from AI technology implementation are truly reflected in corporate profits, if the marginal return rate on capital expenditure declines first, tech stock valuations will come under dual pressure.

Goldman Sachs Research Report Analysis: Momentum Unwinding Shocks Global Stock Markets, AI Spending Boom Conceals Hidden Risks
2026.07.23

AI Impersonating Human Writing Is Polluting the Internet, Substack Decides to Hand Judgment Over to Readers

One scan tells you whether the article is human-written or machine-written.

AI Impersonating Human Writing Is Polluting the Internet, Substack Decides to Hand Judgment Over to Readers
2026.07.23

Podcast Notes | Conversation with Morgan Stanley CEO Jamie Dimon: I'm Not Buying US Stocks or Long-Term Bonds Right Now, Returns on AI Investments Might Differ From Your Expectations

"When I look at AI itself, the money poured into it is enormous. Will there be a return overall? Probably, just like the internet. But will the returns come in the way and timing you expect? Absolutely not."

Podcast Notes |  Conversation with Morgan Stanley CEO Jamie Dimon: I'm Not Buying US Stocks or Long-Term Bonds Right Now, Returns on AI Investments Might Differ From Your Expectations
2026.07.23

Morgan Stanley Research Report Analysis: Software Sector Overly Pessimistic, New Framework Uncovers High-Quality AI Software Targets

The core value of AI should lie in the workflow layer.

Morgan Stanley Research Report Analysis: Software Sector Overly Pessimistic, New Framework Uncovers High-Quality AI Software Targets
2026.07.23

Liang Wenfeng 4-Hour Investor Meeting Transcript

"As long as I can maintain the team's stability, I will definitely achieve AGI. It's just that simple."

Liang Wenfeng 4-Hour Investor Meeting Transcript
2026.07.23

Franklin Templeton: AI Agents Are the Real "Killer App" of Blockchain

Crypto assets may actually be the key to seizing the AI Agent opportunity.

Franklin Templeton: AI Agents Are the Real "Killer App" of Blockchain
2026.07.22

Tonight, Google submits the exam paper for the entire AI market.

The market is not waiting for how much money Google made, but rather the new shipping schedule of Gemini 3.5 Pro.

Tonight, Google submits the exam paper for the entire AI market.
2026.07.22

AI will not gain consciousness, but is becoming society's unconscious

We can do more and more, yet understand less and less. This is a wake-up call for every practitioner who relies on AI tools.

AI will not gain consciousness, but is becoming society's unconscious
2026.07.22

What Jensen Huang Worried About Has Happened: Chinese Open-Source Models Are Taking Over Enterprise AI

This is not a technology competition; it is a reshuffling of AI power.

What Jensen Huang Worried About Has Happened: Chinese Open-Source Models Are Taking Over Enterprise AI
TechFlow Logo

Navigating Web3 tides with focused insights

Contribute An Articleemail
Media Requestsmsg

Risk Disclosure: This website's content is not investment advice and offers no trading guidance or related services. Per regulations from the PBOC and other authorities, users must be aware of virtual currency risks. Contact us / [email protected] ICP License: 琼ICP备2022009338号