Daylila

Information Technology · Monday, 20 July 2026

01 · Briefing · what happened

AI advice made people three times less accurate — and twice as sure

Information Technology 8 min 76 sources

A new study found that access to an AI assistant collapsed people's willingness to say "I don't know" from 44% to 3%, while confidence more than doubled. On the same day, ecologists warned that AI-smoothed photos are seeding false bird sightings into the databases science relies on.

Key takeaways

  • A study found AI assistance cut people's accuracy from 27% to 9% while more than doubling their confidence, and all but destroyed their willingness to say "I don't know."
  • Ecologists warn AI-smoothed photos are seeding false bird sightings into citizen science databases scientists use to track species — only 1,400 of iNaturalist's 610 million images are flagged.
  • TSMC added $100bn to its Arizona plans, taking the total to $265bn, while China's Moonshot AI eyes a listing at a $30bn-plus valuation.

The willingness to say “I don’t know” collapsed

Researchers at three French and Italian universities gave people hard questions, then gave some of them an AI assistant. The group with the assistant got worse at the questions and much more certain about their answers [13].

The numbers are stark. Willingness to say “I don’t know” fell from 44% to 3%. Accuracy fell from 27% to 9%. Confidence rose from 30% to 76% [13].

“People became much worse, the accuracy was only one third, but they were twice as confident,” said Valerio Capraro [13]. He is an associate professor at the University of Milano-Bicocca and one of the study’s authors [13].

The design matters here. The team deliberately picked questions where AI models usually fail — visual details from films, such as the colour of a team’s uniform in Bend It Like Beckham. They used Step 3.5 Flash, a model that was reliably wrong on this set [13]. That was the point: if people had been sensibly handing off to a tool that knew better, the accuracy would have gone up. It went down. Some people who would have answered correctly on their own asked the assistant and became wrong [13].

Paying people to be accurate barely helped. With money on the line, willingness to admit ignorance rose from 3% to 8%, and accuracy from 9% to 16% [13]. Both are still far below the no-AI baselines of 44% and 27%.

The finding echoes work from Wharton earlier this year, which coined “cognitive surrender” for the same effect [13]. People accepted incorrect AI answers 80% of the time, while reporting more confidence than people working alone [13]. The new study sharpens it. The problem is not only that people trust wrong answers. It is that having an answer always available suppresses the habit of noticing you don’t have one.

Capraro said he is particularly concerned about children, who meet these systems before they have built the habit of checking [13]. Common Sense Media, a US nonprofit that rates media and technology for families, this week called the design of Google’s AI search summaries an “unacceptable risk” for students [13].

The angle: if you run a team that has folded AI into research or drafting, the metric worth watching is not output volume. It is how often anyone still says “I’m not sure — let me check.”

The same erosion, in the shared record

A second story, published this morning, shows the same thing happening to a body of evidence rather than a person.

Ecologists are asking birdwatchers to stop running their photos through AI editors. Fake and AI-enhanced images of rare birds are spreading across wildlife photography forums [55]. They are leaking into citizen science platforms such as iNaturalist and Macaulay Library — the databases scientists use to track where species live and how their ranges shift [55].

In a commentary in the journal Nature, researchers reported that hundreds of fake images have already been found on species-recording databases, and said the true scale is unknown [55].

Outright hoaxes are not the main problem. “Nobody is falling for a toucan sighting in Siberia,” said Dr Alexander Lees, an ecologist at Manchester Metropolitan University who wrote the commentary [55]. The damage comes from ordinary tidying. A birder asks an AI tool to remove a branch or make the picture “look better”, and the model quietly fills the gap with features from a different species.

One case: a reported red-winged blackbird in central Brazil, thousands of miles outside its normal North American range. The bird was actually an epaulet oriole, a common local species. The photographer had asked an AI platform to improve the picture, and it added red-winged blackbird parts [55].

The detection gap is the number to hold. Of more than 610 million images on iNaturalist, just 1,400 have been flagged for AI use [55].

“On platforms like ours, regular people are posting information that a scientist could probably never get at scale,” said Tony Iwane [55]. He is iNaturalist’s director of community support and a co-author of the Nature paper. “But the information needs to be accurate” [55].

Apple’s lawsuit lands on OpenAI’s hardware plans

Apple sued OpenAI on 10 July, alleging a pattern of pressing current and former Apple employees to hand over confidential information [7]. OpenAI said it is “not aware of any evidence that this complaint has merit” [7].

The timing is the story. OpenAI is reportedly building its first device — a screenless speaker that can move — and filed confidentially for an IPO in June [7]. A trade secrets case does not need to be won to bite. Discovery, depositions and the risk of an injunction all consume the calendar of the team trying to ship. TechCrunch’s Sean O’Kane argued on the Equity podcast that delay is likely part of the point [7].

One detail worth noting: the complaint leaves out Jony Ive, the former Apple design chief now working with OpenAI on hardware [26]. Naming the most famous person to cross between the two companies would have made the case about him. Apple appears to want it to be about a pattern instead.

Chips: the money keeps moving, more carefully

TSMC, which manufactures chips for most of the industry, is accelerating its Arizona buildout. Chief financial officer Wendell Huang told CNBC the company is committing an additional $100 billion [1]. That takes its total Arizona pipeline to $265 billion. Full-year capital spending rises to between $60bn and $64bn [1].

“We’re seeing this strong-structure, multi-year demand, and we do not plan to leave any food on the table for anybody else,” Huang said [1]. He cited both customer demand and US government support [1]. Not everyone reads it the same way — one analyst told CNBC the scale of the US investment is driven by politics more than business [5].

TSMC is also converting 5-nanometre capacity to the more advanced 3-nanometre node to keep up [1]. The nanometre figure is the rough size of each transistor: smaller transistors mean more of them per chip, and generally more speed for less power.

In China, the money is flowing too, with a wobble. Memory chipmaker CXMT’s $8.6bn Shanghai listing was oversubscribed by institutions roughly 570 times [8]. That is solid, but noticeably cooler than recent Chinese tech IPOs. Reuters put the caution down to a global selloff in chip stocks [8]. Optical component maker Zhongji Innolight rose as much as 8% on Monday, after winning approval for a Hong Kong listing worth up to $8bn [14]. It is likely to be the city’s largest this year [14].

And Moonshot AI, whose Kimi K3 model release last week moved global tech stocks, has told investors it could list within six months. It is closing a round that may value the three-year-old company above $30bn, on annual recurring revenue of $300m — up from $200m in April [3].

The angle: Hong Kong has raised HK$209.9bn across 85 listings in the first half, its strongest in five years, with over 500 applicants queued [14]. Much of that pipeline is AI supply chain. If you track compute costs, that queue is a better leading indicator than any single model launch.

Also moving

Samsung cut jobs across its US display, phone and consumer electronics operations. It said 739 roles in Englewood Cliffs, New Jersey are affected by a headquarters move to Texas, with most staff offered relocation [21]. Around 100 workers at its Plano, Texas office were let go [21]. The New Jersey office opened with fanfare less than a year ago [21]. The split inside Samsung is sharp — its chip division is at record profit while consumer electronics struggles with rising chip costs [21].

Netflix disclosed in a regulatory filing that it paid $587 million in cash for InterPositive [11]. The startup, co-founded by Ben Affleck, makes tools that help fix footage in post-production [11].

Meta users reported Facebook and Instagram outages on Sunday — 4,808 Downdetector reports for Facebook and 2,829 for Instagram in the US, with intermittent access in Singapore. Meta did not immediately comment [10].

The under-covered one: public AI infrastructure

Current AI, a nonprofit founded in February 2025, is trying to build open AI infrastructure that anyone can use without paying a platform. With Bhashini, the Indian government’s AI language division, it built Suno Sutra [17]. It is a pocket-sized device that runs AI in 22 Indian languages with no internet connection, open-sourced for developers to build on [17].

“In India, there are hundreds of different languages and dialects, and right now AI is not representing them,” said chief executive Ayah Bdeir [17]. She joined in January after leading Mozilla’s AI strategy [17]. The nonprofit allocated $3.2 million in grants last month and launched an open-source chatbot at the AI for Good summit in Geneva [17].

It is small money against $265 billion fabs. But it is one of the few efforts aimed at the question of who can use these systems when they cannot pay, and in what language.

02 · Lesson · why it matters

Doubt was doing a job

Not knowing is the alarm that sends you to check, and an always-ready answer switches it off before it can ring.

The number to sit with

Of all the figures in today’s study, the one that matters is not the accuracy drop. It is this: the willingness to say “I don’t know” fell from 44% to 3%.

Accuracy falling from 27% to 9% is bad, but it is the kind of bad we know how to talk about. A tool gave wrong answers, people took them. Fine. The 44 to 3 is stranger. That is not people getting an answer wrong. That is people no longer producing the thought that they might not have one.

And confidence went the other way, from 30% to 76%. Worse and surer, at the same time, in the same heads.

Doubt is a signal, not a gap

We treat not-knowing as an absence — a hole where knowledge should be. It isn’t. It is something the mind actively makes.

That uncomfortable feeling when you half-remember a fact is a working part. It is the alarm that tells you the ground under a claim is thin. It is what makes you check the label, ask the colleague, look it up twice. The discomfort is not a side effect of the mechanism. The discomfort is the mechanism.

Which means it can be switched off without the underlying problem being solved. You can stop feeling uncertain while remaining exactly as uncertain as you were.

That is what the study caught. The participants did not learn the colour of the uniform. They stopped registering that they didn’t know it.

The alarm was calibrated for a slower world

Why would it switch off so easily? Because of what it was measuring.

For most of human life, the cost of resolving a doubt was high. You had to walk to the library, find the person who knew, wait. So the mind evolved a rough rule: raise the alarm loudly, because you will not go to that trouble for a small twinge. The strength of the feeling was tuned to a world where checking was expensive.

Now checking costs one tap. But the alarm was not retuned — it was bypassed. When the answer arrives before the discomfort fully forms, the discomfort never has to do its job, and a signal that never gets used quietly stops being produced.

Notice how thin the safety net is. When the researchers paid people to be accurate, willingness to admit ignorance went from 3% only to 8%. Wanting to be right was not enough. By the time you want to be right, the part of you that would have flagged the problem has already gone quiet.

The same silence, in the record

Today’s second story is the same failure in a body of evidence rather than a person.

A photograph on a species database was never just a picture. It was a check — a piece of evidence another human could dispute. That is what made a public archive of amateur sightings usable by scientists at all. The birder posts, someone else looks, and a wrong identification gets caught.

An AI-smoothed photo is still a photograph. It looks like evidence. But when the model quietly fills a gap with parts of a different species, the picture stops being a check while continuing to look like one. The red-winged blackbird that was really an epaulet oriole did not enter the database as a lie. It entered as a tidied-up photo from someone who wanted a nicer image.

And the database has the same problem the study participants had. Of 610 million images, 1,400 are flagged. That is not a measure of how much is wrong. It is a measure of how little is being noticed. The archive’s own doubt has gone quiet too.

Someone chose that it would always answer

None of this is only about human weakness. There is an arrangement underneath it.

A system that answers every question feels good and gets used. A system that says “I’m not sure” feels broken and gets abandoned. Nobody had to decide to erode anyone’s judgment. They only had to decide, over and over, in a thousand small product meetings, that hesitation looks like failure.

That default is not a fact of the technology. It is a choice, and it poses as a fact. The models can be built to hedge, to show their working, to hold back. Mostly they are built to reply, because a reply is what people rate highly and come back for.

And it genuinely helps. The farmer photographing a dying plant in a language the web barely serves is better off with a system that answers than with silence. Both things are true. The same design that opens the door for her is the one that removes the pause from everyone else.

We inherit each other’s certainty

Here is where the reader is standing, whether or not they use these tools.

You cannot check most of what you rely on. You take the range map, the summary, the confident paragraph. So you do not just inherit other people’s knowledge. You inherit their suppressed doubt, stripped of the flag that would have told you it was thin. An ecologist reads a contaminated database. A student reads a search summary. A manager reads a drafted memo. Each is receiving certainty manufactured somewhere upstream, by someone whose alarm was also quiet.

There is no seat outside this. The researchers who ran the study are inside it. So is anyone who reads a tidy answer and feels that small relief of not having to wonder. That relief used to be the last thing standing between a guess and a fact.

The uncomfortable part is that certainty feels the same either way. From the inside, knowing and being sure are indistinguishable. Which means the honest position on almost everything you believe today is that you do not know which of the two you are holding.

03 · Lab · your turn

The Check You Skip

Answer six questions with an assistant available, and watch how often you stop saying you don't know.

04 · Hope · carry this

Noticing that we had stopped checking is itself an act of checking. The people who caught this were using the very habit they warn is fading, which means it is not gone.

Across the beats