
If you distribute software as a compiled binary, you can no longer assume that withholding the source code will keep outsiders from understanding how it works. A skilled human reverse engineer could always inspect machine code, but the labor cost usually made such projects impractical. AI is now drastically cutting that cost.
As of October 2026, frontier AI agents can operate professional reverse-engineering tools, inspect compiled binaries, recover program behavior, trace call graphs, infer data structures, compare builds, test theories, modify software, and sometimes reconstruct substantial programs without source code. On SRE-Bench, OpenAI reports that GPT-6 Astra fully solved 88.0% of binary challenges on one attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT-5.6 Sol. The benchmark was built from private programs, which reduces the chance that a model could solve the tasks by recognizing source code from training data.
Those headline numbers still need a large asterisk. OpenAI’s cyber evaluation materials describe capability tests run without the production safeguards that apply to normal use, and the strongest SRE-Bench results used a large inference budget across repeated attempts. They do not mean anyone can hand an arbitrary commercial executable to ChatGPT and get the original source tree back five minutes later.
The practical shift is still large. Software distributed only as machine code is becoming easier to understand and modify for people who do not have the source and may not have years of professional reverse-engineering experience.
Game modding already shows what that looks like: developers are reviving abandoned games, repairing mods after vendor updates, reconstructing undocumented formats, and building agent workflows that let an AI investigate an unfamiliar game one subsystem at a time. These same techniques reach much further into proprietary business software, abandoned applications, undocumented protocols, compatibility layers, migration tools, security analysis, and competitive reimplementation.
The cost of writing software is falling. Now the cost of reading software whose source you do not possess is falling too.
Key takeaways
Frontier AI has progressed beyond explaining isolated assembly functions. The best agents can now carry substantial reverse-engineering investigations across many functions, artifacts, tools, experiments, and revisions.
GPT-6 Astra’s near-saturation of SRE-Bench is the strongest benchmark evidence so far, but the result represents a high-compute capability ceiling rather than ordinary consumer access.
Games have become an unusually good proving ground because modifications can often be tested cheaply. Change something, launch the game, inspect the result, then let the agent revise its theory.
Real users have already used AI to remove long-standing game restrictions, repair binary-dependent mods after updates, revive old games on modern platforms, and automate much of the reverse-engineering workflow.
AI still does not reconstruct the original source code in any literal sense. Names, comments, abstractions, types, architecture, and programmer intent may be gone forever. Large, optimized, obfuscated, distributed, or server-dependent systems remain much harder.
The economic consequence could extend well beyond modding. Binary-only distribution becomes a weaker barrier to compatibility work and reimplementation, while verification, domain expertise, proprietary data, online services, distribution, and operational know-how become more important sources of defensibility.
The capability has crossed an important line
AI-assisted reverse engineering is not new. Language models had already been able to explain snippets of assembly, guess what functions do, translate decompiler output, and help write Ghidra scripts for years.
The crucial change is agency.
A modern coding agent does not have to stare at an entire 100 MB executable as text. It can operate more like a junior reverse engineer sitting in front of a workstation.
It asks a reverse-engineering tool to decompile one function. It follows the callers. It searches for strings. It examines cross-references. It identifies a likely structure. It renames functions. It records an inference. It examines another subsystem. It writes a small script to test an assumption. It compares two versions of a binary. It runs the program and observes the result. When reality contradicts its theory, a good agent revises the theory and continues.
That loop is the important capability.
Ghidra already provides disassembly, decompilation, graphing, scripting, and other compiled-code analysis tools. Projects are now wrapping those capabilities in interfaces designed specifically for AI agents. Cellebrite Labs’ ghidra-rpc, for example, lets an agent decompile functions, trace callers and callees, rename symbols, define data structures, patch code, and diff binary versions. Similar agent integrations exist for IDA Pro.
That architecture solves one of the biggest problems with using LLMs for reverse engineering: an executable can contain tens or hundreds of thousands of functions, far more information than should be dumped into a model context at once.
The reverse-engineering database becomes the agent’s external memory.
Instead of remembering the whole program, the model can progressively turn an anonymous sea of FUN_14001A230-style functions into a partially labeled model of the software. That persistent state also changes how errors can be handled. A weak hypothesis does not have to live forever in chat history. The agent can rename a function, attach evidence, revise a type, compare a later observation against an earlier annotation, and update the project when the theory stops fitting.
That persistence is easy to underestimate. Long reverse-engineering jobs fail when the investigator forgets why a symbol was renamed, which assumption produced a patch, or which observation contradicted an earlier guess. External project state gives the model somewhere to keep those decisions where later tool calls can inspect them.
That is much closer to how serious human reverse engineering works.
What “reverse-engineering a binary” actually means
People may be tempted into imagining a compiler running backward, but compilation is not generally reversible that way.
Suppose the original developer wrote something like a class called InventoryManager, with methods bearing descriptive names, useful comments, several source files, custom types, and carefully chosen abstractions. After compilation, especially in an optimized release build with debugging symbols removed, much of that information does not need to remain in the executable.
The CPU needs machine instructions. It does not need the programmer’s specific labels and comments.
A disassembler translates machine-code bytes into assembly instructions. A decompiler goes further and tries to reconstruct higher-level control flow and expressions. Good decompilers can produce pseudocode that looks remarkably like C, but it is an interpretation of the machine code rather than recovered original source.
AI adds another semantic layer.
A model may notice that an anonymous function accepts an object pointer, checks a character identifier, calls another function associated with a game world, then returns a Boolean. From surrounding evidence, it may infer that the function validates whether a character belongs in a particular level.
That inference can be extremely useful even though the model has never recovered the developer’s original function name.
This creates several different levels of success.
▪ At the easiest end is recognition. What does this function appear to do?
▪ Then comes mapping. Which functions implement this subsystem, which data structures do they manipulate, and how are they connected?
▪ Then modification. Can the agent locate the behavior controlling a feature and change it without breaking unrelated behavior?
▪ Then reconstruction. Can it recreate enough of the subsystem in fresh source code to reproduce its behavior?
▪ At the far end sits whole-program reimplementation, where an agent builds new software that behaves like the original even though it never had the source.
Frontier systems can already do meaningful work at every one of those levels. Reliability falls as scope and ambiguity increase.
SRE-Bench shows how quickly the frontier moved
SRE-Bench was created specifically because ordinary coding benchmarks are poor tests of reverse engineering.
Its researchers spent more than 5,000 expert hours building 19 private programs averaging about 16,900 lines of code, then applied 44 anti-analysis techniques to create 262 binary instances and 1,572 deterministically graded tasks. The programs cover areas including network protocols, firmware, games, file formats, and malware-related analysis. Because the programs were created privately for the benchmark, the design specifically reduces the contamination problem that makes many software benchmarks easier to game through memorization.
The original paper looked much less dramatic than today’s leaderboard.
GPT-5.6 Sol, then the strongest evaluated system, achieved a 61.4% per-instance score but fully solved only 31.5% of the instances. Claude Fable 5.1 fully solved 26.9%. The revised paper says those results show that strong source-code security performance does not automatically transfer to binary analysis.
Then GPT-6 Astra arrived. OpenAI reported 88.0% full solves on a single attempt and 99.2% pass@4, while GPT-5.6 Sol managed 55.9% and 68.7% under the same comparison. That is an unusually large capability jump over the earlier results, and it pushed SRE-Bench close to saturation almost immediately.
It also makes SRE-Bench less useful as a long-term frontier test unless harder successors arrive. A benchmark that moved from “largely unsolved” to nearly saturated almost immediately tells us something important about the rate of progress, even if it can no longer separate future models very well.
Popular AI’s earlier look at GPT-6 Astra versus GPT-5.6 Sol found the same broader pattern in software engineering: Astra’s strongest case is long, tool-driven work where an agent repeatedly investigates, changes something, checks the result, and continues.
Reverse engineering fits that profile almost perfectly.
More on GPT-6 Astra:
Different reverse-engineering tasks produce very different model rankings
SRE-Bench should not be interpreted as “GPT-6 Astra can reverse engineer 99.2% of software.”
The task distribution, agent harness, tools, reasoning budget, protections, and grading criteria all influence the result.
A July 2026 ARIMLABS benchmark asked seven frontier models to extract concrete indicators from real malware samples through static reverse engineering. Claude Opus 5 led with 87% recall, followed by GPT-5.6 Sol at 82%, Kimi K3 at 80%, and GLM-5.2 at 74%. Plaintext indicator recovery was becoming close to commoditized, while the harder compiled samples still separated the models sharply.
SentinelLABS tested something different again.
It reconstructed an eight-stage professional investigation of the fast16 sabotage implant and asked models to keep the reverse-engineering project trustworthy as later evidence invalidated earlier conclusions. GPT-5.6 Sol was the only publicly available model in that evaluation to complete the full eight-stage investigation. Other frontier systems produced useful local analysis but could not maintain the same project-wide coherence through the whole sequence.
The researchers’ most interesting observation was not that Sol never made mistakes. It did.
The difference was that it could sometimes repair the investigation after making them.
It could withdraw a bad conclusion, determine which downstream artifacts depended on it, revise those artifacts, rerun checks, and continue.
That is a very different skill from explaining one function correctly.
Senior human reverse engineers were still needed to define quality, challenge weak conclusions, expose blind spots, and approve final findings. SentinelLABS describes the best current role as supervised investigative agency rather than autonomous replacement.
That description applies surprisingly well outside malware analysis too.
Why games are becoming the perfect AI reverse-engineering playground
Game modding has several properties that suit coding agents.
Most PC games place a substantial executable system on hardware the user controls. Games also tend to have strong behavioral or visual feedback. If you change character validation and the previously forbidden character loads into the level, you have learned something. If the game crashes, you have learned something else.
An agent can keep iterating.
That makes many game-modding problems cheaply verifiable or “hill-climbable.” You do not necessarily need a perfect mental model of the entire program before every experiment. You need a hypothesis, an intervention, and a useful oracle telling you whether the hypothesis survived contact with reality.
Games also contain a wonderful mess of different reverse-engineering targets: native C++, managed .NET assemblies, Lua scripts, custom archives, shaders, serialized assets, proprietary sprite formats, physics behavior, save files, network protocols, and engine-specific data structures.
One game can therefore exercise nearly every part of an AI agent’s software-analysis toolchain.
And unlike a traditional reverse engineer who may spend days learning an unfamiliar engine before doing useful work, an agent arrives with broad prior knowledge of common compilers, engines, file formats, APIs, assembly idioms, mod loaders, and debugging techniques.
The expert knowledge has been compressed into the model. The target-specific knowledge can be learned during the session.

Claude Code helped break a 13-year-old Disney Infinity restriction
One of the cleanest recent examples is InfinityUnlocked, a mod for Disney Infinity 1.0.
Disney Infinity normally restricts characters to their corresponding playsets. The mod allows characters to cross those boundaries.
The project’s author explicitly credits Anthropic’s Claude Code with Opus 4.6 for the reverse-engineering work. The restriction turned out to be enforced across several layers of native code rather than through one obvious gate. The finished mod combines 17 executable patches with changes to three data files.
The technically interesting part is how the solution was found.
The project reports that Claude helped trace the relevant call graph, identify repeated validation sites across different game systems, understand the surrounding x86 logic, and work through approaches that initially failed.
This is the kind of problem older coding assistants struggled with. There was no repository full of nicely named C++ classes. The useful abstraction existed only after somebody reconstructed it from the executable.
The model helped turn anonymous machine code back into a workable theory of the game’s rules.
The result is also a reminder that reverse engineering does not require rebuilding the whole application. A model that can accurately map the small slice of a huge binary that controls one behavior may already be enough to enable sophisticated mods.
That is probably the most important near-term use pattern.
AI helped revive Chromatron from old binaries
Piotr Migdał’s revival of the 20-year-old puzzle game Chromatron illustrates a much larger-scale attempt.
Migdał began without prior binary reverse-engineering experience. He connected Claude Code to Ghidra and tried to reconstruct the game from old Windows XP and PowerPC binaries so it could run on modern Apple Silicon and the web.
The process required repeated decompilation, comparison, correction, and hand-holding.
Early versions required considerable hand-holding. The model reproduced some structural elements but invented details, including fonts and text that were not present in the decompiled evidence. Other decompilation routes recovered positioning more accurately but struggled with assets. The project tried several combinations of models and tools before converging.
That failure behavior is as informative as the success.
LLMs are extraordinarily good at filling gaps with plausible code. Reverse engineering punishes that instinct.
If the agent is reconstructing a twenty-year-old game’s renderer and encounters an unexplained constant, “something that would make sense here” is not the answer. The binary is the answer.
Successful AI-assisted reverse engineering therefore depends on making evidence authoritative. Decompiled instructions, runtime observations, file contents, screenshots, test results, and binary diffs need to overrule the model’s intuition.
The more an agent can be forced into that evidence loop, the better it becomes.
AI is already repairing mods after commercial game updates
Another common reverse-engineering chore is version migration.
A mod may depend on internal functions at particular locations in a game’s executable. The developer updates the game, compiler output changes, addresses move, and the mod stops working.
Ken Ng documented using Claude Code to repair FFVIIHook after a June 2026 Final Fantasy VII Rebirth update broke the mod. He had little previous decompiler experience and used Ghidra to compare old and new versions of the executable while Claude handled much of the scanning and scripting.
Some functions were straightforward to relocate. One difficult function had hundreds of call sites and resisted the obvious matching strategy. The eventual process used repeated candidates and real crash feedback to eliminate wrong answers.
When an official updated build later appeared, Ng compared it with the AI-assisted result. The mod worked, and the principal version-specific values matched, although the official build contained additional details the AI-assisted patch had missed.
That example captures the current frontier unusually well.
The AI did enough reverse engineering to produce useful software for a real commercial game while still missing parts of the implementation.
“Point Claude at any game” is now an actual project
At the end of September 2026, the universal-modder project pushed this idea a step further by packaging game modding as a reusable AI-agent workflow.
Its pitch is almost absurdly direct: point an AI coding agent at a PC game, let it work out the engine and modding route, inspect the relevant code or data, build a mod, test the result in the running game, and write down what it learned for the next agent.
The project includes playbooks for a wide range of engines and technologies. Its reverse-engineering path can connect agents to tools such as Ghidra, IDA, ILSpy, Cpp2IL, RenderDoc, and other analysis systems. It also restricts its stated workflow to owned, offline or single-player games and refuses anti-cheat or DRM bypasses.
The examples are already more ambitious than changing a texture.
The project’s knowledge base records an Age of Empires II: Definitive Edition project that created a new civilization and involved reconstructing the game’s .sld sprite format. It also records a Terraria experiment that decompiled boss behavior into a simulator for an AI-agent test. Those examples sit alongside dozens of game-specific notes, which is useful evidence for the broader point: the workflow is being packaged for reuse rather than treated as a one-off demo.
The more important development is workflow reuse across many mods.
Once an agent has learned how one engine stores sprites, how one loader expects mods to be packaged, or how one executable exposes a certain class of game system, that knowledge can be written into skills and handed to the next agent.
Reverse-engineering expertise becomes software.
Why targeted modifications are much easier than recovering an entire program
The phrase “AI can reverse engineer binaries” covers a huge range of difficulty.
A targeted modification can be surprisingly tractable. A modder may already know the visible behavior they want to change. Strings, nearby functions, runtime effects, old versions, community documentation, file names, or mod APIs give the agent starting points.
The problem might reduce to locating one subsystem among thousands.
Whole-program understanding is different.
A large native game may contain millions of instructions spread across rendering, physics, scripting, input, animation, serialization, networking, audio, platform code, third-party libraries, and engine internals. The binary can mix compiler-generated scaffolding with handwritten logic. Templates, inlining, link-time optimization, and aggressive compiler transformations can make one source-level concept appear in many places or make several concepts disappear into one optimized function.
The model can inspect all of it eventually. Maintaining a correct global theory is harder.
This is why persistent annotations, tests, call graphs, automated scripts, version control, and structured investigation records are so important. A long-running agent needs external state for much the same reason a human reverse-engineering team does.
Managed software is a different proposition from optimized native code
Not all binaries throw away the same amount of information.
A managed .NET or Java program may preserve class names, method signatures, metadata, and other high-level structure. Unless heavily obfuscated or ahead-of-time compiled, decompilation can sometimes recover source-like output that is strikingly readable.
Native C and C++ binaries are generally harder. Optimized Rust and Go introduce their own recognizable compiler patterns and runtime structures, but they are still compiled machine code.
Debug symbols can make the job dramatically easier. Stripping those symbols makes it harder. Obfuscators, virtualization-based protection, packers, encrypted assets, self-modifying code, integrity checks, and custom loaders increase difficulty again.
Architecture matters too.
A familiar x86-64 Windows executable backed by decades of tooling is a friendlier target than obscure firmware running on an unusual processor with poorly documented peripherals.
AI does not erase these differences. It reduces the amount of specialist labor needed to navigate them.
The model does not need perfect decompilation to be useful
This point is easy to miss.
A human reverse engineer rarely needs to recreate every original source file before accomplishing something useful. Neither does an AI.
Suppose a proprietary image editor stores project files in an undocumented format.
A developer who wants an importer may only need to determine how the relevant chunks, indexes, metadata, compression schemes, and checksums work. Reconstructing the application’s brush engine or licensing UI provides no value.
Likewise, a game modder interested in NPC behavior does not need to reverse engineer the renderer.
A compatibility project may only care about an undocumented protocol.
A preservation project may need the save format and game-state logic.
An enterprise migration may care about extracting records from a dead application that no longer runs on supported operating systems.
AI’s ability to search enormous code surfaces and identify the small portion related to a concrete goal is therefore potentially more valuable than perfect source reconstruction.
Black-box reimplementation makes the proprietary-software question even bigger
Binary analysis is only half of this story.
Sometimes an AI may not need to inspect the binary internally at all.
The 2026 MirrorCode benchmark asks coding agents to reproduce complete software programs without source-code access. The agent can run a reference executable, read documentation, supply inputs, inspect outputs, and construct its own independent implementation.
Across 25 target programs, the strongest evaluated model configuration scored 56%. Models successfully rebuilt substantial software, including a roughly 16,000-line bioinformatics toolkit. Some large attempts consumed enormous inference budgets, including one run reported at roughly $2,600 over 19 days.
This is black-box reverse engineering rather than decompilation.
It may eventually matter even more commercially.
A proprietary application’s implementation can be hidden while its behavior is necessarily exposed to users. If an agent can interrogate the product systematically, generate thousands of experiments, infer edge cases, and write a replacement that passes behavioral tests, access to the original implementation becomes less essential.
The commercial software moat does not disappear. But hiding source code becomes a weaker moat by itself.
What frontier AI could do with proprietary software
There is a large legitimate territory between “I have the source” and “I am trying to steal somebody’s product.”
Consider an organization with a business-critical application from 2004 whose vendor disappeared. The company still has the executable, databases, and documents but cannot compile or maintain the program.
An AI reverse-engineering agent could help document the file formats, reconstruct business rules, identify dependencies, extract data, recreate integrations, or port essential behavior into maintained software.
An acquisition team could analyze a legitimately obtained executable where source provenance is incomplete.
▪ A developer could build an interoperable importer for a proprietary format.
▪ A platform developer could understand the interface required to support old plug-ins.
▪ A security team could audit a commercial binary whose source is unavailable.
▪ A preservation group could reconstruct enough of an abandoned program to keep historical files usable.
▪ A software vendor could recover behavior from its own binaries after source loss.
▪ A hardware community could document an abandoned device protocol.
And a competitor could potentially use lawful clean-room techniques to recreate compatible behavior without incorporating the original code, although the legal boundaries depend heavily on jurisdiction, contract, purpose, and exactly what is copied.
AI makes each of these projects cheaper because a small team can explore more hypotheses in parallel.
Proprietary code is not the same thing as proprietary capability
The implications vary dramatically depending on where the valuable part of a product actually lives.
A purely local desktop application exposes a great deal to its owner. The executable has to contain enough information for the machine to perform the work.
A cloud application is different.
If the desktop program is only a client that authenticates against a remote service, the customer does not possess the server implementation. Reverse engineering the client may reveal the API calls, data structures, validation assumptions, UI behavior, or protocol, but it does not magically reveal the remote code, databases, models, internal operations, or infrastructure.
This pushes an old software-security lesson into much broader commercial relevance: anything that must execute on the user’s machine should increasingly be treated as inspectable.
That was already technically true for sufficiently skilled reverse engineers.
The difference is the potential number of people who can perform the inspection and the speed with which they can do it.
Companies whose defensibility consists mainly of “the interesting algorithm is hidden inside our executable” should pay attention.
Companies whose value comes from proprietary datasets, constantly changing online services, network effects, customer relationships, operational expertise, regulated access, or large physical infrastructure are in a very different position.
AI reverse engineering attacks obscurity more directly than it attacks those other moats.
Version diffing could become one of the most practical enterprise uses
Software vendors reveal information whenever they release an update.
A reverse engineer can compare version A and version B and ask what changed.
Humans already do this for vulnerability research, malware analysis, compatibility work, mod maintenance, and undocumented API discovery. Agentic tooling makes the process easier to automate. Cellebrite’s current Ghidra tooling explicitly includes binary version matching and changed-function analysis.
This has benign applications everywhere.
A company can investigate why a vendor update broke its integration. A mod developer can find relocated functions. A security team can understand which parts of a closed application changed after a security patch. A migration team can determine whether a proprietary file format changed.
Diffing is also much easier than understanding a binary from zero because one build becomes evidence about the other.
AI is very good when it can ask a constrained question such as: which previously known function became this function?
The remaining limitations are substantial
The current capability deserves attention precisely because there is no need to exaggerate it.
▪ Models still invent explanations
Chromatron’s fabricated fonts and text are a useful warning. Reverse engineering creates endless opportunities for a model to mistake plausibility for evidence.
A convincing explanation of assembly is not proof that the explanation is correct.
▪ Large programs create coherence problems
SentinelLABS found models capable of good local analysis that nevertheless failed to maintain a correct investigation across stages. The problem was not always intelligence at the function level. It was project management, state, correction, and closure.
This is exactly what becomes difficult in a large game or commercial application.
▪ Optimization destroys human-friendly structure
Functions get inlined. Loops change shape. Variables disappear into registers. Template instantiations multiply. Dead code vanishes. Compiler-generated helpers appear everywhere.
A decompiler is reconstructing one possible readable representation from the surviving machine behavior.
▪ Runtime behavior can carry the real answer
Static analysis may tell you what code could do. Sometimes you need to see what it actually does with real files, real hardware, real network responses, or specific program state.
That means debuggers, traces, logs, captures, emulation, instrumentation, and controlled experiments remain important.
▪ Remote services place a hard boundary around what the binary reveals
A client cannot reveal server-side source code it never received.
It can reveal what it sends, what it expects back, and what assumptions it makes about those responses.
That can still be valuable. It is not the same as possessing the server.
▪ Obfuscation and anti-analysis still cost the agent time
SRE-Bench deliberately included anti-analysis mechanisms because ordinary binaries and protected binaries are not equivalent.
Protection therefore still has economic value.
The question increasingly becomes whether it creates a durable barrier or merely increases the amount of inference the attacker must buy.
▪ Long jobs can be expensive
MirrorCode demonstrates that difficult reimplementation tasks can consume days and enormous token budgets.
Humans are no longer the only scarce resource. Compute becomes another one.
▪ Provider safeguards can restrict access
OpenAI classifies GPT-6 Astra at the Critical cyber capability level and applies stronger safeguards around advanced cyber use. Some of its most aggressive capability evaluations were run without the normal production safeguards. The theoretical ceiling of a hosted frontier model is therefore different from what an ordinary user can access for every target and objective.
That creates another reason local and open models will remain strategically interesting even when they lag the frontier.
Hosted reverse engineering also creates a confidentiality problem
Giving an AI agent a proprietary binary means giving it access to proprietary material.
If the model runs remotely, organizations should understand what gets uploaded, retained, logged, used for training, exposed to administrators, or governed by enterprise contracts before pointing the agent at confidential software.
This is the same issue companies already face when coding agents touch private repositories, except binaries can be even easier to overlook because people instinctively think of an executable as something meant to be distributed.
A vendor’s public executable is one thing. An unreleased firmware image, internal build, licensed SDK component, customer application, or software received under NDA is another.
Popular AI has previously examined the broader risk of letting hosted coding agents touch private repositories. Users who need to keep code and binaries off third-party servers can also explore local AI workflows, although today’s local models generally trade frontier capability for more control.
The agent itself should also be isolated when examining unknown code. Reverse engineering often rewards an agent for opening odd files and experimenting with suspicious artifacts. Popular AI’s coverage of the Friendly Fire coding-agent exploit shows why unrestricted execution of material being analyzed is a bad default.
A model investigating an unknown binary should not need your personal SSH keys to do its job.
More on sandboxing agentic AI:
Reverse engineering proprietary software is not automatically illegal
The legal picture is much more specific than either “reverse engineering is illegal” or “if you own a copy, anything goes.”
In the United States, two important Ninth Circuit decisions established meaningful room for reverse engineering aimed at discovering unprotected functional elements.
In Sega v. Accolade, the Ninth Circuit held under the facts before it that disassembly was fair use where it was the only practical way to access unprotected functional elements and the copier had a legitimate reason for doing so. Accolade’s purpose was compatibility with the Sega Genesis.
In Sony v. Connectix, the Ninth Circuit held that intermediate copying during reverse engineering of the PlayStation BIOS was fair use in the circumstances of that case, allowing Connectix to develop its emulator.
Section 1201(f) of the U.S. Copyright Act also contains a reverse-engineering provision that, under specified conditions, permits circumvention when necessary to identify and analyze elements needed for interoperability of an independently created program.
The European Union’s Software Directive contains a related decompilation provision. Article 6 allows decompilation without authorization from the rights holder when it is indispensable to obtain information necessary for interoperability of an independently created program, subject to explicit conditions and restrictions on how the resulting information can be used.
None of those rules is a universal permission slip.
Copyright in expressive code, contractual terms, trade-secret obligations, anti-circumvention law, unauthorized access law, patents, and redistribution of copyrighted game assets can create different problems.
U.S. contract law adds another complication. In Bowers v. Baystate, the Federal Circuit held, applying the law involved in that case, that a license restriction against reverse engineering could support a separate contract claim even where copyright law did not provide equivalent protection. The treatment of those clauses can vary with the jurisdiction and facts.
So the safe conclusion is boring but accurate: purpose, jurisdiction, how the copy was acquired, license terms, technological protection measures, what information is extracted, and what is distributed afterward can all change the answer.
AI changes the cost of the technical work. It does not repeal the legal rules around that work.
Software vendors should assume client-side code can be understood
For software engineering, one implication is immediate.
Secrets should not be protected merely by hoping nobody can understand the executable.
Hard-coded credentials were already a bad idea. Client-side authorization decisions were already weaker than server-side enforcement. Proprietary protocols were already discoverable with enough effort.
AI increases the number of people for whom “enough effort” is affordable.
Every application need not move to the cloud in response. Moving computation to a server creates privacy, availability, cost, account-dependency, and user-control tradeoffs of its own.
Engineers should instead be explicit about the trust boundary.
A client program can enforce local UX rules. It should not be treated as a trustworthy guardian of secrets against the machine’s owner.
This is old security advice arriving in a new economic environment.
Old software becomes much more valuable
There is another side to the story.
Humanity owns an enormous amount of software whose maintainability has effectively expired.
Source code has disappeared. Build systems no longer work. Dependencies vanished. The original company died. Nobody understands the file format. The application only runs under an obsolete operating system. A government department has a critical tool written by people who retired fifteen years ago.
Historically, “we have the executable but not a maintainable source tree” could turn a useful system into technical archaeology.
Frontier agents are unusually well suited to archaeology.
They can combine decompilation with knowledge of old APIs, forgotten compilers, operating systems, file formats, and modern replacement libraries. They can write compatibility shims, tests, documentation, parsers, converters, and replacement components as they discover what the old system expects.
The Chromatron project is a charming game-sized example of something that could become economically serious.
Reverse engineering may become part of ordinary legacy modernization.
Tests become more important when the code did not originate with humans
MirrorCode points toward a related change in how software specifications work.
If an agent can repeatedly interrogate a reference program and construct another program that passes the same behavioral tests, then the observable contract becomes more important than the original implementation.
That makes comprehensive testing more valuable.
A poor test suite tells an AI replacement very little about the obscure cases humans care about.
A rich collection of conformance tests, historical inputs, golden outputs, protocol traces, invariants, and failure cases can act as a machine-readable description of what the old software actually means.
That changes the value of old test artifacts. A forgotten regression corpus or directory of customer files may describe behavior more precisely than a stale design document. When an agent is recreating a system from observation, every verified input-output pair constrains the replacement. Edge cases that once looked like annoying test maintenance can become part of the specification.
The same principle applies when AI writes new systems. Fast generation increases the value of independent checks because plausible code is cheap and trustworthy behavior is not. A generated replacement that passes the obvious happy path may still mishandle malformed files, old versions, locale differences, timing behavior, or obscure state transitions. The article’s reverse-engineering examples keep returning to the same practical rule: evidence has to outrank fluent explanation.
As implementation becomes cheaper, proving that an implementation satisfies the intended behavior consumes a larger share of engineering attention.
Software engineering is moving from creation toward supervision and verification
There is already empirical evidence of this shift outside reverse engineering.
A 2026 longitudinal study of professional software engineers found that 82% reported spending less time writing code when using AI coding assistants. The researchers also observed a broader shift from creation toward verification and described supervisory engineering work centered on directing, evaluating, and correcting AI output.
Anthropic reached a related conclusion from a privacy-preserving analysis of roughly 400,000 Claude Code sessions. People made most of the planning decisions while Claude made most execution decisions. Domain expertise also predicted better outcomes and more work completed per instruction.
That is almost a description of the emerging AI reverse engineer.
The human says: this restriction should not apply here, find the real enforcement path and prove the modification works.
The agent does the mechanical excavation.
The human still needs to know what constitutes convincing proof.
Reading binaries changes software development beyond security
Coding agents were initially framed as code generators.
Reverse engineering makes them something broader: program-understanding machines.
That changes the useful scope of software automation.
An agent can potentially enter a project through source code, documentation, a running application, a binary, an API, a network capture, an old database, screenshots, or test cases and start reconstructing how the system behaves.
Greenfield coding is only one entry point.
That creates a much larger body of software work for agents. Maintenance, migration, integration, compatibility engineering, porting, modding, preservation, and documentation of undocumented systems all become more automatable.
The software industry has spent decades accumulating programs that nobody wants to touch. An agent that is unusually patient about reading ugly machine-generated pseudocode may be exactly what that pile of technical debt needed.
The developer economy changes when implementation is no longer the scarce ingredient
This development strengthens a trend already visible in AI coding.
Implementation skill is not disappearing, but the market value of implementation in isolation is under pressure.
Anthropic’s study also found that people outside traditional software occupations were already succeeding at code-producing Claude Code sessions at rates close to software professionals on average, while task-specific expertise remained associated with better outcomes.
Reverse engineering pushes the same phenomenon in a different direction.
A knowledgeable Disney Infinity fan who understands exactly what the game should permit may now be able to investigate its native executable with an agent despite lacking years of x86 reverse-engineering experience.
▪ A biologist may be able to recreate an abandoned research utility.
▪ An archivist may be able to recover data from obsolete software.
▪ A business analyst who understands an old internal application’s rules may be able to supervise its migration.
▪ A modder with a strong idea may become capable of modifications that previously required a specialist reverse engineer.
The ability to describe the problem, recognize incorrect behavior, and define a successful result becomes more valuable relative to manually producing every instruction needed to get there.
The expert’s role shifts upward. One expert can supervise more mechanical investigation while keeping responsibility for the judgment that decides whether the evidence is good enough.
SentinelLABS reached essentially this conclusion in malware analysis. The scarcity of skilled reverse engineers has historically limited how much software can be analyzed. A senior analyst supervising capable agents can expand the amount of investigative work that one expert can oversee.

Closed-source software loses one part of its historical moat
For decades, distributing binaries without source imposed a large information cost.
That cost never made reverse engineering impossible. It made it expensive enough that most users did not bother.
AI reduces that cost.
The commercial effect will vary by product.
A desktop utility whose only defensibility is a clever local algorithm faces a different future from a cloud service with proprietary real-time data and a huge operational backend.
A mature engineering package with decades of validation, customer trust, support, documentation, certification, specialized domain knowledge, and plug-ins still has enormous value even if somebody can understand portions of its executable.
But a small proprietary utility that charges heavily mainly because nobody else knows its undocumented format may find that position harder to defend.
Compatibility can become cheaper. Migration can become cheaper. Replacement can become cheaper.
And “there is no public documentation” stops being as decisive an obstacle.
It will also produce mountains of questionable code
Cheaper implementation does not automatically produce better software.
AI-generated pull requests already illustrate the basic economic problem: generation costs can fall faster than verification costs. Popular AI has covered how AI-generated contributions can move work from authors to maintainers.
AI reverse engineering can create the same imbalance.
A thousand people may soon be capable of creating binary patches.
That does not mean a thousand people are capable of proving those patches are safe.
Mods can corrupt saves. Patches can behave differently under another build. Decompiled logic can be misunderstood. An agent can accidentally modify an adjacent code path. A reimplementation can pass visible tests while failing on obscure real-world inputs.
Software provenance therefore becomes more important, not less.
So do reproducible tests, signed releases, isolated execution, version checks, rollback paths, code review, and clear documentation of what an agent actually changed.
The limiting factor moves downstream.
More on vibe code slop:
The most interesting change may be that existing software becomes programmable again
Source code is usually treated as the editable form of software and a binary as the finished artifact.
That division is becoming less absolute.
The binary remains more difficult to modify than good source code. It lacks much of the information developers intentionally put into a source tree. It may be optimized, stripped, protected, and architecturally opaque.
Yet the gap between “I possess this executable” and “I can meaningfully change what this program does” is shrinking.
▪ For games, that means mods nobody bothered making because the reverse-engineering cost was too high.
▪ For abandoned software, it means another chance at maintenance.
▪ For businesses, it means legacy systems can sometimes be understood even after institutional knowledge disappears.
▪ For software vendors, it means client-side obscurity offers less protection.
▪ For developers, it means AI coding increasingly includes understanding existing systems rather than merely generating new ones.
And for the developer economy, it moves scarcity toward the parts machines still cannot cheaply infer: domain knowledge, product judgment, validation, distribution, trust, proprietary data, operational execution, and the ability to decide what software ought to do.
AI reverse engineering makes binaries a weaker barrier
The frontier in late 2026 is no longer “AI can help a reverse engineer.”
The stronger statement is now supportable: frontier AI can perform substantial reverse-engineering work itself when given the right tools, enough compute, a concrete objective, and a way to verify progress.
GPT-6 Astra’s near-saturation of SRE-Bench is the clearest quantitative evidence in this article. Disney Infinity, Chromatron, Final Fantasy VII Rebirth, universal-modder, malware research, and other real projects show what the capability looks like outside a benchmark.
The capability remains fallible. Models lose architectural context, invent explanations, and benefit enormously from expert supervision. Whole-program reconstruction is still much harder than finding and changing one subsystem.
Even with those limits, binaries are becoming a weaker informational barrier. Modding feels the change first because modders are willing to experiment and can often verify results cheaply. Legacy maintenance follows because organizations have a direct economic reason to recover behavior from software they can no longer build. Proprietary software, interoperability, migration, security analysis, and commercial reimplementation are next in line.
For developers, the practical lesson is straightforward: treat client-side code as inspectable, preserve tests and behavioral evidence, isolate agents that analyze unknown binaries, and keep human review focused on the claims that are hardest to verify.
The developer who can write software with an AI agent is already becoming ordinary. The next step is a developer who can point an agent at software nobody has the source for and start working anyway.
Explore more from Popular AI:
Start here | Local AI | Builds & gear | Autonomy & policy | Fixes & guides | Popular AI podcast
















How much do you think AI-powered reverse engineering will change the value of keeping software source code proprietary?