Agentic Reverse Engineering: Bridging Binary Disassembly and Model Context Protocol with REA

Reverse engineering native binaries and compiled software has traditionally required hours of manual effort inside disassemblers. Security engineers alternate between reading assembly instructions, writing Frida scripts for dynamic memory tracing, and correlating offsets across decompiled function graphs. When LLMs entered developer workflows, early attempts at binary analysis relied on copy-pasting raw assembly snippets into prompt buffers. This approach failed quickly due to context limits, lost execution state, and missing cross-references.

Background

The open-source project REA (Reverse Engineer Anything) changes how autonomous security agents interact with binary targets. Built as a Model Context Protocol (MCP) server running on Node.js 22+, REA bridges static analysis tools like Ghidra and Hopper directly with agentic runtimes. Rather than feeding unparsed disassembler text into an LLM context window, REA presents disassemblers, process debuggers, and AST parsers as executable MCP tools.

In standard reverse engineering tasks, an analyst inspects relocation tables, traces system calls, and verifies control-flow graphs (CFGs) to identify entry points. REA automates this workflow by exposing granular disassembler actions to the model context. An agent can query symbol tables, inspect dynamic allocations, or retrieve basic blocks from function graphs without drowning in unparsed hex dumps or thousands of lines of assembly.

By defining explicit MCP schemas for binary analysis tools, REA transforms raw disassembler APIs into structured inputs. Agents invoke operations such as listing cross-references, fetching decompiled C pseudocode for a specific address, or monitoring runtime network calls while maintaining state across iterative queries.

Autonomous tools should not replace binary analysis judgment; they should reduce the mechanical friction of navigation and offset calculation.

Challenges

Building an agentic interface for native binaries introduces major technical obstacles:

  • Context Explosion: Raw assembly dumps for medium-sized functions can burn 50,000 tokens in a single response frame. REA solves this by truncating disassembler outputs to localized basic blocks and target function signatures.
  • Address Space Layout Randomization (ASLR): Static disassembly addresses in Ghidra or Hopper rarely match dynamic memory addresses during runtime execution. REA translates static relocation offsets to live memory base addresses during dynamic process tracing.
  • Tool State Synchronization: Coordinating headless Ghidra scripts, Hopper disassembler bridges, and dynamic tracing hooks without corrupting target memory or locking file handles requires a central orchestration daemon.

When analyzing Position-Independent Executables (PIE), an agent must query binary structures incrementally. If an agent requests a full section decompilation at once, the context window fills with assembly noise, causing reasoning drift and hallucinated instruction pointers. REA enforces narrow tool call parameters, forcing the agent to request only relevant basic blocks and symbol metadata.

Results & Lessons

Setting up REA in a security testing environment requires a single terminal command. The interactive setup script detects local disassemblers, configures agent runtimes, and writes validated MCP client definitions:

npx rea-agents setup

Once initialized, the generated MCP configuration connects your preferred coding agent directly to the REA execution engine:

{
  "mcpServers": {
    "rea": {
      "command": "npx",
      "args": ["-y", "rea-agents", "mcp"]
    }
  }
}

You can also install the REA skills package directly into agentic development environments:

npx skills add morluto/rea --skill reverse-engineer-anything

Our evaluation revealed three key takeaways when deploying agent-assisted disassembly:

  • Targeted Symbol Querying: Restricting agent tools to function-level decompilation prevents hallucinated memory addresses and retains token budget for complex control flow logic.
  • Deterministic Bridge Tracing: Using fixed MCP schemas to communicate with Ghidra and Hopper eliminates syntax errors that occur when models generate raw disassembler scripts on the fly.
  • Human Verification Boundaries: Agents excel at mapping API calls and finding candidate vulnerabilities, but manual validation remains essential before executing binary patches or proof-of-concept exploits.

Combining structured Model Context Protocol tools with binary analysis engines marks a pragmatic step forward for security engineering. By isolating disassembler navigation inside deterministic MCP tool calls, security teams can automate routine binary analysis tasks while maintaining clear boundaries around agent execution.

Press Cmd K to search برای جستجوی سایت از Cmd+K استفاده کنید