Draft poster · the finished cut will appear here
What happened
Chang Liu, Edward Raff, and Kristopher Micinski argue in an arXiv paper that a decompiler can produce clean, buildable C while quietly changing what the original binary does. In their evaluation, some reconstructions passed the supplied tests but diverged on additional inputs. Across eight systems and nine configurations, 4.9 percent of passing candidates diverged, with one system reaching 13 percent.
How it works
The authors introduce Decompile-Diverge as a second check. It generates a driver for the reference function, grows a fuzzing corpus from that reference, and sends the same inputs to the reconstruction. The outputs can then be compared without treating the shipped test suite as a complete behavioral specification.
The paper also compares build success with behavioral agreement. On 300 GitHub library functions and 287 CVE-grounded functions, a refinement LLM raised Ghidra-derived build success from 75 to 90 percent. The matched-behavior rate fell from 74 to 62 percent, and up to one tenth of the disclosed vulnerability cases showed Crash Absence in the reconstructed code.
Why it matters
The useful distinction is between source that looks plausible and a program that still behaves like the original. For AI-assisted ports, a green compiler and inherited tests are evidence of progress, not proof of preservation. A comparison loop that attacks the boundary between the reference and the reconstruction belongs beside ordinary unit tests.
What we don’t know yet
This is an arXiv preprint, not a universal failure rate for LLM decompilers. The percentages depend on the systems, datasets, fuzzing oracle, and the authors’ definitions of a match and a crash.
Source: arXiv, Chang Liu, Edward Raff, and Kristopher Micinski, September 4, 2026 - original paper