Daylila

Biotech & Longevity · Sunday, 23 August 2026

01 · Briefing · what happened

The first blind test of AI antibody design, and the week biology checked its own tools

Biotech & Longevity 2 min 14 sources

Twenty-nine organisations committed 511 AI-designed antibodies before a lab made a single one. On the ranking task all but one model lost to picking clones at random - and days later a sleuth found 54 ageing papers built on an antibody that binds the wrong organism.

511

AI-designed antibodies tested

committed by 29 organisations before any were built [1]

39%

of random picks beat the free option

against 9.8-13.8% for the AI submissions [1]

46.4%

of design-from-scratch entries failed

did not bind at all, or bound but could not be made into a drug [1]

54

ageing papers, wrong antibody

each lists one raised against an E. coli protein to hunt a human one [2]

At a glance

  • Twenty-nine organisations committed 511 AI-designed antibodies before a lab made a single one of them. [1]
  • On the ranking task, only 9.8% to 13.8% of AI picks beat the lab's free shortcut - take whichever version of the antibody turns up most often. [1]
  • Picking versions at random beat that same shortcut 39% of the time, so most models did worse than chance. [1]
  • Every winning algorithm but one, from Washington University, also came in below random picking. [1]
  • On the task the models suited best - strengthening an antibody a lab already had - about 13% of entries improved it twenty-fold or more. [1]
  • Days later, a sleuth found at least 54 ageing papers naming an antibody raised against a bacterial protein, not the human one they wanted. [2]
  • Nature Medicine put the 'biological age' clocks through the same kind of test, across 51 studies and 16 clocks. [5][6]
  • Two studies found medicine aimed at the wrong places: antibiotic overuse concentrated in rich countries [7], and long-term illnesses taking 2.5% of health aid while causing 59.5% of the world's disease burden [14].

Forces in play

Flattering self-scoring High

Nearly every antibody-design benchmark had been run on results that already existed, which makes almost any method look good [1].

Independent checking Building

AIntibody is the field's first blind, commit-first test, modelled on the protein-structure contest that made AlphaFold's case [1].

Unchecked lab tools High

Testing an antibody before an experiment is slow and costly, so it usually is not done; one sleuth found 54 ageing papers naming the wrong one [2].

AI that does deliver Steady

An Essex team used protein-design software to rebuild 672 antibodies to stay stable inside cells, and is releasing them free [3].

In play The AIntibody consortium — ran the blind contest, and also reported its own results [1] Washington University — the only entrant that beat random picking on ranking, 50% against 39% [1] Aureka — matched the lab's best hand-made antibody on the strengthening task [1] Xencor — won the design task on the published rules, with a molecule that would likely fail development [1] Sholto David — the independent biologist who found the 54 mismatched papers [2]

How it unfolded

  1. 2024 the challenge is announced, its rules published in advance [1]
  2. Early 2025 29 organisations commit 511 sequences; no strengths are revealed to them [1]
  3. Tue a Palo Alto startup unveils a virtual-cell model meant to predict how a whole cell reacts [4]
  4. Wed the results publish: on ranking, all but one model lose to random picking [1]
  5. Wed the FDA clears a rare bone-disease drug on a blinded, placebo-controlled trial of 63 adults [8]
  6. Fri Nature reports 54 ageing papers listing an antibody raised against the wrong organism [2]

Where this points

Watch whether a second round runs with an outside body holding the answers - the organisers say this one rested on their own integrity, and closing that gap is what would turn a one-off into a standard [1].

Full briefing

Why the score changed when the order changed

An antibody is a protein that grips one target and holds on. A better one grips harder without falling apart. Computers now propose them, and for years those proposals were judged the cheap way: run the method over results a lab measured long ago, count the agreements. The expensive way is to build every proposed molecule and measure it. Almost nobody had done that [1].

AIntibody changed the order of events. Twenty-nine organisations committed 511 sequences first. Only then did a lab make them and measure them [1].

The order is the whole difference. Score a method on results that already exist and the answer was in reach while the method was being built. It sits in the training data, in the published literature, in a hundred small tuning choices. None of that is fraud. It is simply not a test.

Two features of this problem make the flattery worse. Antibody sequence is not a language: at most positions the commonest building block is the inherited one, and inherited is usually not the version that grips hardest [1]. A model that learns what is typical learns the wrong target. And the real training data barely exists in public - discovery results sit inside companies, bound by patents and rarely machine-readable [1].

The paper’s own baselines disagree, and it says so. Earlier published work put the share of randomly picked clones that beat the standard control at about 30%. Measured directly in these three clusters, it was 39% [1]. Both figures are in the same paper. Most submissions were below either one.

The check has weak points too, and the organisers name them. One target. A consortium that both ran the contest and reported it. Blinding that rested on their own integrity rather than a technical lock [1]. A first blind test beats none, but it is not yet the structure-prediction contest it was modelled on.

The week’s other finding is why this reaches past one field. Aled Edwards, a biochemist in Toronto, told Nature that the ready-made molecules biologists buy to find things are routinely used without being checked, because checking is slow and expensive [2]. He expects it to worsen as models read the literature to pick disease targets. A paper naming the wrong antibody hands that error to every model that reads it, invisibly.

Two contrasts landed the same week. The FDA cleared a second drug for a rare bone disease on a randomised, blinded, placebo-controlled study of 63 adults [8][9][10]. It cleared a gene therapy early on a stand-in measure, with confirming evidence still years away [11]. And the president named Heidi Overton to run the agency that decides which of those counts [12][13].

02 · Lesson · why it matters

The test you can pass with the answers in front of you

A prediction only counts as evidence if it was locked in while the answer was still out of reach.

How it works

  1. A method is scored on results that already exist
  2. The answer was in reach while the method was built
  3. So the score measures fit, not foresight
  4. Make everyone commit before the answer exists
  5. The same methods drop below picking at random

The twist

A prediction is only evidence if it was locked in while the answer was still out of reach - otherwise you are measuring how well something fits what it has already seen.

Where you've seen this

Investing

a trading rule tuned on the last ten years of prices always wins the last ten years

Weather models

one fitted to past storms can replay them perfectly and still miss next week

Your own memory

you keep the hunches that came true and quietly drop the ones that did not

Exams

a paper stops measuring anything the moment the questions leak

The catch

A blind test is slow, costly, and only ever asks what its organisers chose to ask - this one used a single well-studied target, so a method could fail here and still be useful elsewhere.

Full lesson

The moment the order changed

For years, the papers said computer-designed antibodies were working. The evidence was real. A team would take their method, run it across antibody results some lab had measured and published, and count how often the method picked the winners. It picked them often.

Then twenty-nine organisations were asked to do it the other way round. Commit first. Hand over your sequences before anyone makes them. Only afterwards does a lab build every one and measure how tightly it grips.

On one of the three tasks, all but one method fell below what you would get by closing your eyes and pointing.

Nothing about the methods had changed. Only the order.

Why the answer leaks backwards

When you score a method against results that already exist, the answer is loose in the world while the method is being built.

It is in the training data, because those published results are exactly what everyone trains on. It is in the literature, which the people building the method have read. It is in a hundred small choices made along the way: which measure to use, which cases to exclude, which version to keep when two versions disagree. Each choice was made by someone who already knew, roughly, what the right answer looked like.

None of that is cheating. Nobody looked up the answers. The answers were simply in the room.

The word for the fix is old and dull: blind, and before the fact. You state what you predict, you seal it, and then the world is asked. It is why a drug trial hides which patients got the pill. It is why a sealed envelope in a magic trick only impresses if it was sealed first.

Why this arrangement is normal

Backward-looking tests are cheap. You need a computer and a published dataset. Forward-looking ones need somebody to synthesise five hundred molecules and run them through a machine, three times each, for months.

So the cheap test became the standard test, and every method that survived was one that did well on the cheap test. Nobody decided this. It fell out of who pays.

That is worth noticing carefully, because it does not look like a decision at all. It looks like how the field simply is. Most arrangements do. Somebody bore a cost, somebody could not, and the shape that resulted got called normal.

The same week, the tools underneath

While that result published, an independent biologist reported on at least fifty-four papers about cell ageing. Each had listed an antibody raised against a bacterial protein, then used it to hunt a human one.

Antibodies are what biologists use to find things. Checking that one binds what the label says is slow and expensive, so it is often skipped. And the cost of skipping it does not land on the lab that skipped it. It lands on everyone who reads the paper afterwards - now including the models that read papers by the million to decide which proteins are worth aiming a drug at.

An error that nobody paid to catch is copied faster than it was made.

Where you are in this

You are downstream of all of it. The medicines that reach you are the ones a field decided were worth the next hundred million, and it decided using tests like these.

But you also run the cheap version, constantly. You remember the hunch that came good and lose the three that did not. You judge a decision by how it turned out, knowing the outcome, and feel sure you would have called it. Nobody hands you a sealed envelope for your own life.

The people who ran this blind test know it too. They said so in their own paper. One target. A consortium that both organised the contest and reported the score. Blinding that rested on their own honesty rather than a lock. They built the best check anyone has built here, and they still had to name what it cannot see.

That is roughly where all of us stand. Not outside the system, holding the marking scheme - inside it, having marked our own work for years without noticing, and only now finding out by how much.

03 · Lab · your turn

Commit Before You Look

Rehearse choosing which molecules to build when the only evidence about your predictor came from answers it had already seen.

04 · Hope · carry this

Twenty-nine groups handed over their best work before anyone could grade it, and the unflattering answer got published anyway. Fields become trustworthy one uncomfortable test at a time.

Across the beats