TruaceTracing the truth around AIWednesday, August 5, 2026
TRV-2026-0487Certified recordPeer-reviewed

The Illusion of Thinking

Recent generations of frontier language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answers. While these models demonstrate improved performance on reasoning benchmarks, their fundamental capabilities, scaling properties, and limitations remain insufficiently understood. Current evaluations primarily focus on established mathematical and coding benchmarks, emphasizing final answer accuracy. However, this evaluation paradigm often suffers from da…

Science · The Trace — both readings · certified 2026-07-22 · v1 · article view · machine-readable

Current reading — gain

Large Reasoning Models generate detailed thinking traces before answering and demonstrate improved performance on reasoning benchmarks, with advantage over standard LLMs on medium-complexity controllable puzzles.

Current reading — problem

Frontier Large Reasoning Models face a complete accuracy collapse beyond certain puzzle complexities and exhibit a counterintuitive scaling limit where reasoning effort declines despite adequate token budget.

What this doesn’t fix

Findings are bounded to controllable puzzle environments and reveal that LRMs have limitations in exact computation and inconsistent reasoning across puzzles.

Evidence

Reader signal

How should this claim be treated?

Cite this record

Truvace Impact Record TRV-2026-0487, v1: “The Illusion of Thinking.” Truvace, 2026-07-22. /record/TRV-2026-0487 (accessed at citation time). sha256 a56d29a25f21da80

Calibration history

Every change to this record since certification, in the open. None yet — the reading has held since it entered the record.

  1. Certifiedv1a56d29a25f21

    Certified into the record

Verify this record
How to verify without trusting this page

Fetch the canonical text of any version from /api/record/TRV-2026-0487 and hash it yourself — for example shasum -a 256 on the saved canonical field. The result must equal content_hash, and each version’s text ends with prev:followed by the prior version’s hash (version 1 chains to 64 zeros). If a single character of any version had been altered since certification, the chain would not reproduce.