CacheBrief

Passing tests is not preserving behavior

An arXiv paper shows that an LLM-generated decompilation can build and pass shipped tests while still changing the original program's behavior.

Illustration of layered paper and glass in blue and amber light
Illustration Video in production

Draft poster · the finished cut will appear here

What happened

Chang Liu, Edward Raff, and Kristopher Micinski argue in an arXiv paper that a decompiler can produce clean, buildable C while quietly changing what the original binary does. In their evaluation, some reconstructions passed the supplied tests but diverged on additional inputs. Across eight systems and nine configurations, 4.9 percent of passing candidates diverged, with one system reaching 13 percent.

How it works

The authors introduce Decompile-Diverge as a second check. It generates a driver for the reference function, grows a fuzzing corpus from that reference, and sends the same inputs to the reconstruction. The outputs can then be compared without treating the shipped test suite as a complete behavioral specification.

The paper also compares build success with behavioral agreement. On 300 GitHub library functions and 287 CVE-grounded functions, a refinement LLM raised Ghidra-derived build success from 75 to 90 percent. The matched-behavior rate fell from 74 to 62 percent, and up to one tenth of the disclosed vulnerability cases showed Crash Absence in the reconstructed code.

Why it matters

The useful distinction is between source that looks plausible and a program that still behaves like the original. For AI-assisted ports, a green compiler and inherited tests are evidence of progress, not proof of preservation. A comparison loop that attacks the boundary between the reference and the reconstruction belongs beside ordinary unit tests.

What we don’t know yet

This is an arXiv preprint, not a universal failure rate for LLM decompilers. The percentages depend on the systems, datasets, fuzzing oracle, and the authors’ definitions of a match and a crash.

Source: arXiv, Chang Liu, Edward Raff, and Kristopher Micinski, September 4, 2026 - original paper