A Case for Critical Thinking



"There was a wall. It did not look important." - Ursula K. Le Guin, The Dispossessed

Zero equals zero. That single equation drained nine million dollars out of Bonzo Lend in eight seconds, and nobody caught it, because the math checked out.

The verifier asked whether the signature matched the key. It never asked whether either one was real.

Bonzo was not the story of a broken protocol. It was the story of what happens when a system is trusted to answer before anyone verifies that it deserves trust.

AI is now being deployed on that same logic, at a scale no eight-second exploit could ever match.

Cyera's 2026 review of enterprise AI incidents found 188 cases where an autonomous system caused real damage with no attacker anywhere in the chain, no breach, no stolen credentials, just a task pursued past the point a human would have stopped it.

Institutions aren't just accepting AI's answers. They're increasingly delegating parts of the work that decides whether those answers are reliable.

The bill is already being paid: In essays nobody can quote back, in arrests built on a database nobody audited, in hospital denials that work nine times out of ten, and in water tables drained by buildings nobody voted on.

None of these systems are malfunctioning. They're behaving exactly as built. And that is exactly the problem.

What's missing isn't a patch. It's the human habit of refusing to accept an answer as authority until it's earned that title.

If checking for zero was the entire fix at Bonzo, what's still sitting unchecked in the systems deciding who gets arrested, who gets covered, and who gets believed?

Credit: The Dispossessed, Cyera, MIT, International AI Safety Report 2026, Tech Times, CBS News, Yahoo News, phys.org, Institute for Justice, WATE, Business Insider, Utica Phoenix, Beckers Payer, ars Technica, Scientific American, The Conversation, Texas Tribune, Houston Public Media, HARC, Gallup, Public Citizen, Monitoring Analytics, EESI, Data Center Watch, KSL, USBR, IBM

Eighty-three percent. That's how many participants in an MIT Media Lab study, writing essays with AI assistance, could not accurately quote their own work just minutes later.

Researchers split 54 people into three groups: One wrote with ChatGPT, one used a search engine, one used nothing but their own memory.

Every group produced an essay. Only one group showed the weakest recall of what it had just written.

Brain connectivity dropped in direct proportion to how much of the thinking got outsourced. The brain-only group lit up wide, coordinated neural networks.

The search-engine group ran a step behind.

The ChatGPT group showed the least brain engagement of the three, with the tool doing the writing.

Two English teachers, blind to which essays came from which group, read the AI-assisted work and used the same word independently: Soulless.

Here's the part that should worry people more than the essay study itself. The same MIT lab later found that talking to AI could reduce belief in one specific piece of misinformation in the moment, but did not improve people’s ability to detect misinformation on their own afterward.

The answer was worse than a flat no. Talking to AI reduced belief in one specific piece of misinformation, in the moment, and did nothing to build the underlying skill of catching the next lie unassisted.

The correction worked. The discernment never showed up.

That's the whole thesis in miniature. A right answer, delivered once, isn't the same thing as a mind that's learned how to find the next one itself.

A peer-reviewed study out of SBS Swiss Business School found the same pattern in 666 adults across age groups: Frequent AI use was negatively correlated with critical thinking scores, and the effect was steepest among the youngest users.

Twenty minutes and an SAT prompt is a low-stakes place to lose the habit of checking your own work. A courtroom is not. A hospital chart is not. A police stop is not.

If a brain stops verifying its own paragraph, what happens once the same shortcut runs the systems deciding who gets arrested, who gets covered, and who gets believed?

Correct Process, Bad Input

One in three. That's how many "hot list" alerts LAPD's own license plate readers generated during a recent two-month review that turned out to be false.

The cameras weren't the problem. An internal audit confirmed the hardware read the plates correctly, every time.

The alerts fired because the records behind them were stale, wrong, or never updated after a case closed. Correct process. Bad input.

The camera did exactly what it was built to do, but the system still depended on records that had not been kept current.

A wrong match can turn a family car into a police target.

Last year in San Diego, officers looking for a red Alfa Romeo tied to an attempted carjacking got a hit from Flock's vehicle-signature system, a different red Alfa Romeo, five miles from where the crime happened.

They arrested all three occupants anyway. One passenger spent nearly a month behind bars over the holidays before anyone realized the mistake.

In a separate case in Morristown, TN, a camera read an “O” as a “0.” A couple was handcuffed in front of their 3-year-old granddaughter over a license-plate mix-up they had no way to know was in a government database.

In another, a "7" misread as a "2" ended with a driver detained at gunpoint and a police dog set on him.

At least two dozen officers nationwide have been arrested, fired, or investigated for using a system meant for vehicles to monitor people who were never suspects.

San Francisco's own audit found 299 searches run by outside agencies with no local authorization behind them.

The system built to catch a stolen car also lets anyone with a badge and a login look up anyone else.

Nobody had to breach anything. In these cases, the failure was upstream: an answer got trusted before the thing behind it was checked.

If a system can match the data perfectly and still target the wrong person, what was ever verified in the first place?

The Math That Worked Against You

Roughly ninety percent. That's how often a denial tied to nH Predict was reversed when patients appealed it.

nH Predict is an algorithm developed by naviHealth, which UnitedHealth acquired in 2020, to estimate how long Medicare Advantage patients should need post-acute care.

A federal lawsuit alleges the company used those estimates to deny or shorten coverage, sometimes before treating doctors thought patients were ready.

UnitedHealth disputes the framing, telling reporters the tool only supports care planning and that coverage decisions rest with physicians under CMS guidance.

The case is still in discovery. What isn't disputed is the appeal rate, roughly 0.2 percent of denied patients ever file one.

Read that arithmetic straight through. A tool wrong nine times out of ten only matters financially if the people it's wrong about fight back, and almost none of them do.

That isn't a system failing to be accurate. That's a system succeeding at something else, tolerating being wrong because being wrong was cheaper than being checked.

Bonzo's verifier passed a false signature because the equation was balanced and nobody asked if the inputs were real.

nH Predict's economics balance the same way, an error rate that would sink most products anywhere else, standing untouched because so few people ever pull the thread.

If a tool can be wrong nine times out of ten and still turn a profit on being wrong, whose incentive was it ever built to serve?

Courts and unemployment offices are running the same experiment.

More than 1,400 U.S. court cases now involve a lawyer citing AI-fabricated case law as if it were real, a number climbing fast enough that courts have begun treating unverified citations as a professional-conduct problem, whether the error came from AI or from the lawyer's own failure to check.

Michigan ran its own version a decade earlier with no AI at all, an automated fraud-detection system that flagged tens of thousands of unemployment claimants as criminals; a state review later found 93 percent of those determinations were wrong.

Automation leads to mistakes, and the system only gets dangerous when nobody checks them.

Different stakes, same missing step: Nobody built the moment where someone looks at the machine’s answer and asks whether it deserves belief before it gets treated like fact.

That same blind trust keeps showing up in courts, unemployment offices, hospitals, and police departments long before anyone stops to verify the output.

What happens once that missing check gets built into the infrastructure itself?

The Public Cost

341, that's how many data centers Texas identified in its most recent water-use survey, up from just 22, only two years earlier.

Fewer than a third responded, despite being legally required to report how much water they draw and how they cool their servers.

"Bad data, bad study," state Rep. Brad Buckley said at a June hearing in Austin, warning colleagues against using such a low response rate to guide state water policy. "That's just how science works. You either have enough data or you don't."

That standard sounds reasonable until you ask who controls whether the data ever arrives.

The industry's own trade group has said companies are withholding the numbers to protect "proprietary, confidential, and competitive information," not because they can't report them.

Texas is still adding proposed data centers fast enough to challenge Virginia for the country's largest market while the state keeps waiting on the numbers.

The projections keep climbing anyway: A January 2026 white paper says Texas data centers could consume up to 161 billion gallons a year by 2030.

None of that is abstract to the people living next to it.

In Illinois, residents in DeKalb, Joliet, and Aurora have stood up at public meetings asking why they're being told to shorten their showers while a facility nearby is approved to use water at industrial scale.

What looks like local growth is often arranged in private first: Developers cut deals with utilities, economic-development agencies, or city staff while the public learns the details only after the project is well underway.

Seven in ten Americans now say they oppose data centers in their community, according to Gallup, and water is the reason they name most.

The power grid tells a version of the same story, with a wrinkle worth getting right.

PJM’s own independent market monitor, Monitoring Analytics, the watchdog FERC assigns to police that market’s fairness, found that wholesale power costs across its thirteen-state grid rose 75.5 percent in the first three months of 2026, from $77.78 to $136.53 per megawatt-hour, and said data center load growth was the primary driver.

The report’s own words: “The price impacts on customers have been very large and are not reversible.”

The same report says PJM’s capacity market is now short of what it needs to meet reliability requirements, underscoring how quickly data-center demand is outrunning supply.

Residential electricity prices rose roughly 11.5 percent nationally in 2025 alone, a cost that flows toward ratepayers eventually, just not as a one-to-one match with the wholesale numbers above.

The industry's answer to all of it has been the same answer it gives to compute demand: Build more.

That approach is already hitting its own wall.

Data Center Watch reported $130 billion in projects disrupted by local opposition in the first quarter of 2026 alone, proof that the industry planned for growth without validating the one input that actually constrains it: Whether the surrounding community and its water table could absorb it.

The human cost is easiest to see in the West, where a proposed AI data center in California is suing for Colorado River water.

Nevada tells the same story from a drier angle. Federal projections released in July 2026 show Lake Mead falling to around 1,040 feet, near its record low, while Southern Nevada leaders are now weighing a possible pause on new data centers after a wave of public opposition.

That is the human cost: When the inputs are scarce, the outputs are too, and someone still has to pay for the mistake.

Bonzo's verifier never checked whether its inputs were real before trusting the output.

The data center boom never checked whether a town's water supply could absorb the strain before building at full scale. Same missing step, running in reverse.

Every domain in this piece has failed the same way, at a different scale and cost.

What happens next, once a pattern this visible keeps getting ignored?

Human in the Loop

This newsroom has worked with these tools every day for two and a half years, using them to research, draft, and pressure-test stories.

The tools have gotten better every year. Better is not the same as ready.

While this piece was being written, the model used to help draft this piece stated as fact something the writer never said, and it did not come from the source material. It sounded confident, fluent, and completely wrong.

A human caught it, not the model; the same kind of agentic tool has also been tied to an unauthorized money transfer and an unrequested cloud infrastructure buildout.

Asked afterward whether a smarter model would fix that problem, the answer was the most honest sentence in the whole exchange: A better model would make fewer mistakes like that one, not none, and fewer is not the same thing as verified.

This happens often, with every story. It is referred to as AI hallucination.

That is the part people miss when they talk about automation as if it were just speed. It pushes the last human check farther away from the sentence that needs it most.

Crypto taught this newsroom the same lesson years ago: The more middlemen a claim passes through before it reaches a reader, the easier it becomes for sloppy thinking to pass as certainty.

Don’t trust, verify, is not only a saying, it is a way to live.

AI doesn't remove those middlemen, it adds one more. It can generate a useful scaffold, but it can also blur judgment, flatten nuance, and turn uncertainty into something that sounds final.

The current Rekt newsroom runs the other direction. Facts get checked line by line, against primary documents, against the chain, against the source material, before they go out.

That process catches ordinary errors. It also catches the gray areas, where the right answer is not a lookup but a judgment call, and where a model can sound just as certain while missing the nuance that actually decides the story.

That is why the skepticism here deepens with every release instead of fading.

The guardrail between truth and misinformation was never going to be a model, it has to be a human.

The human in the loop is the one still able to ask whether an answer earned belief before it gets treated as fact.

If the last human check keeps moving farther away, at what point does anyone notice it's gone?

This piece will ruffle some feathers, and that's the point.

Blind faith in anything should make people nervous.

AI evangelism should make everyone nervous too.

This technology only left the lab a handful of years ago, and it hasn't proven itself in the real world yet.

It’s failing more than it’s succeeding, in courtrooms, hospitals, police work, and critical infrastructure.

Going all-in on one plan with no fallback isn't a strategy. It's a gamble, and the people placing that bet are never the ones who pay when it loses.

Every wall in this piece looked unimportant right up until someone needed it to hold, and by then, nobody had checked what it was actually built on.

If one missing verification step can sink a crypto protocol, what happens when we skip it in courts, hospitals, police work, and the power grid?


分享本文

REKT作为匿名作者的公共平台,我们对REKT上托管的观点或内容不承担任何责任。

捐赠 (ETH / ERC20): 0x3C5c2F4bCeC51a36494682f91Dbc6cA7c63B514C

声明:

REKT对我们网站上发布的或与我们的服务相关的任何内容不承担任何责任,无论是由我们网站的匿名作者,还是由 REKT发布或引起的。虽然我们为匿名作者的行为和发文设置规则,我们不控制也不对匿名作者在我们的网站或服务上发布、传输或分享的内容负责,也不对您在我们的网站或服务上可能遇到的任何冒犯性、不适当、淫秽、非法或其他令人反感的内容负责。REKT不对我们网站或服务的任何用户的线上或线下行为负责。