Live process instrumentation · 15 languages · native and managed

Your AI agents can finally
see native code.

Java, Python, Node and .NET are the easy four. flux-dbg attaches to C, C++, Rust and Go — and eleven more — inside a process that is already running, with no agent installed before the incident started. Capture what it is doing, prove what went wrong, leave it running.

Because "see live production" is only true if something was watching before it broke.

sonic-demo.cast — recorded, not simulated

live attach to a SONiC switch · 49s · real capture

The bug reproduces in production and nowhere else. Attach to the process that is misbehaving, capture the arguments it is actually being called with, and detach — without a redeploy, and without losing the state you were trying to look at.

The incident is live and the logs don't have the field you need. Add the instrumentation now, on the running service, with an expiry and a rate breaker enforced inside the target so the probe cannot become the second incident.

The failure is on a customer's box, running a build you shipped months ago, and "please restart it and try again" is not an acceptable sentence. Attach to the daemon in place, collect an evidence bundle, and leave the process running. The TAC workflow →

An agent reading logs reasons about the past. 44 MCP tools let it place a probe where its hypothesis says the answer is, watch the code run, and revise — the loop a good engineer runs, at machine speed, inside limits it cannot widen. See a real run →

Attached to, and still running

nginxPostgreSQLRedisetcd containerdCockroachDBTiKVFRR PyTorchNumPyCeleryFastify LangChain.jsGraalVM
15
languages, native and managed, one attach model
0
restarts, redeploys or launch flags required
36/36
stripped nginx workers probed live, 60/60 requests captured
CET
shadow-stack-armed targets probed live, verified on bare metal

The problem

Production is the one place you cannot look

The bug only exists there

It needs their traffic, their config, their peers, their timing. The local reproduction that would let you debug it properly is the thing you cannot build — which is why the ticket has been open for three weeks.

Logging is a guess made in advance

Every log line was a prediction about which failure would happen. The field you need now is the one nobody predicted, and adding it means a build, a review, a deploy, and waiting for the failure to come back.

The tools that could look need permission you don't have

A debugger stops the process. A profiler was not loaded. An agent had to be there before the incident started. Each answer arrives one restart too late, and the restart destroys the state.

All three are the same problem wearing different clothes: you have to decide what to observe before you know what went wrong. flux-dbg removes that ordering constraint — you decide afterwards, on the process that is already running.

Three components

One binary, three ways to see inside a program

What you can do with it

Questions you can answer without a deploy

Debug live production

Attach to the process that is misbehaving right now, capture the arguments it is actually being called with, and detach. The state you wanted to look at is still there, because nothing restarted.

Trace a binary you didn't build

No source tree, no debug package, no symbol table. Probe sites come from exported dynamic symbols or DWARF where it exists. The nginx on the proof page is a stock stripped distro binary.

Replay a run deterministically

Record every instruction, syscall, thread switch and signal to one file, then replay it as many times as you like and get the same execution each time. GDB attaches to the replay, with reverse execution on small recordings — see the limits on Proof.

Find where the latency went

Span-kind probes fold into a per-function digest — call count, p50, p95, max, with sample arguments — so "which call is slow" is one command, not an afternoon.

Capture evidence for an incident

Arguments, return values, struct fields, strings, globals and multi-hop dereferences, as NDJSON. Fault-guarded: a bad capture emits null rather than taking the process with it.

Probe a hardened binary

Intel CET shadow stacks reject exactly the transfer a naive trampoline performs. fr trace detects an armed shadow stack and emits a CET-legal sequence instead — verified on silicon that enforces it.

Intel CET →

Let an agent investigate

The same operations over MCP, with an expiry and a per-site rate breaker applied by default — because the caller is often autonomous and won't remember to clean up.

Coverage

Native and managed, one tool

The agent-model products cover four managed runtimes. flux-dbg covers fifteen languages, because it instruments the binary rather than the bytecode — which is also why compiled languages are in the list at all.

CC++ RustGo JavaKotlin ScalaGroovy ClojurePython Node.jsTypeScript Bun RubyPHP Erlang

Amber are compiled-to-native. Each language has a live-attach end-to-end test in the repository — and that includes GraalVM native-image, where Java is compiled ahead-of-time to a native binary with no JVM left to instrument. Full matrix, including what isn't supported →

Safety

A probe should never become the incident

Limits live inside the target, not in the CLI. Kill the attaching process, lose the network, crash the collector — the expiry and the rate breaker still fire, because they are counters in the injected library.

Expiry

--ttl-sec N — probes go inert N seconds after they load. For the attach nobody remembers.

Rate breaker

--max-hz N — a single site over N hits/sec is disabled, per site. For the probe on a per-packet path.

Limits only tighten

If the service owner set a ceiling and an attach asks for more, the tighter one wins. An agent cannot widen it.

The guardrail model →

Read this before adopting

What isn't built yet

This is a single-engineer project competing with funded products. Here is what they have and it doesn't, kept on the front page rather than discovered after you've committed.

  • No PII redaction and no blocklists. A capture emits whatever is at the address. A probe on an auth path will capture tokens.
  • No conditional capture. Sampling exists; only when user_id == 42 does not.
  • No .NET / C#. The one runtime the agent-model products cover and this doesn't.
  • No enterprise control plane. No RBAC, SSO, SCIM, audit log or agent pools.

The first two are the current priority. Full comparison →

Capability · Probe Injector

It's already running.
That used to be
the problem.

Compiled normally, started normally, nobody prepared it for this — and that is exactly the process you can attach to. Probes go in while it serves traffic, capture what it is really being called with, and come out the same way.

The difference that decides everything

Datadog Dynamic Instrumentation, Rookout and Lightrun are good at what they do, and their instrumentation genuinely is dynamic — once their agent is loaded. The catch is how it gets loaded.

RuntimeHow an agent-model debugger gets inWhen
JVMjava -agentpath:/…/agent.so -jar app.jarprocess start
Pythonimport lightrun — a source changeprocess start
.NETCORECLR_ENABLE_PROFILING=1 + profiler pathprocess start
flux-dbgfr trace attach <pid>any time

So the real question is never "can it debug production". It is "was the agent already there when production broke". If it wasn't, the answer is a restart — and a restart destroys the state you were trying to look at.

Sources: vendor installation docs captured 2026-09-10. Their configuration changes carry the same shape — "To apply the changes, the application must be restarted."

The lifecycle

Four commands, and the process never stops

01 · ATTACH

Agent attach into a live pid. No ptrace stop of the whole process, no dlopen in the hot path.

02 · FIRE

Captures arguments, locals, struct fields, strings, globals, multi-hop dereferences. Fault-guarded — a bad capture emits null.

03 · GUARD

The probe expires on its own and disables any site that runs away, enforced inside the target.

04 · DETACH

--remove-all restores the original bytes. The process keeps running, unmodified.

Mechanism, not magic

What this actually does to your process

The claim reads like marketing, so here is the mechanism and the failure modes. If these aren't satisfying, nothing else on this site matters.

BEFORE — THE FUNCTION AS COMPILED 55 48 89 e5 41 54 push rbp; mov rbp,rsp; push r12 …rest of the function… 5 bytes replaced, atomically AFTER — PATCHED, PROCESS STILL RUNNING e9 ?? ?? ?? ?? jmp trampoline 41 54 (displaced) …rest, untouched… trampoline: capture() → run displaced bytes → jmp back

Every other instruction in the process is the one the compiler emitted. There is no code cache and no dispatcher, so an unprobed path costs nothing at all — and --remove-all writes those five bytes back.

What gets written

A 5-byte jmp rel32 at the probe site, pointing at a trampoline that calls the capture and returns control at a natural instruction boundary. No DBT engine, no code cache, no dispatcher — unprobed paths run 100% natively.

How it's written safely

The cross-modifying-code sequence from Intel SDM 8.1.3: flip byte 0 to 0xCC, membarrier to serialize every CPU's pipeline, write bytes 1–4 while every thread traps at byte 0, barrier again, flip byte 0 to 0xE9. A SIGTRAP net catches any thread inside the window. No INT3 remains.

If the tooling dies

The limits are counters inside the injected library, not policy in a server. Kill the CLI, lose the network, crash the collector — the expiry and the breaker still fire.

When the hardware fights back

If the target has an Intel CET shadow stack armed, the ordinary relocated-call sequence is ROP-shaped and the CPU kills the process. The injector detects this at runtime and emits a CET-legal sequence instead. How that works →

If it isn't sure

A build-ID mismatch between the spec and the running binary is refused outright rather than patched hopefully. That check fired unplanned while recording the demo on the proof page.

The questions people actually ask

Before you point this at anything you care about

Can it crash my process?

Yes, in principle — anything that writes to another process's memory can. Being specific is more useful than reassurance: the patch is applied with the cross-modifying-code sequence from Intel SDM 8.1.3, captures are fault-guarded so a bad read emits null rather than a signal, and a build-ID mismatch is refused outright instead of patched hopefully.

The honest evidence is that a dedicated fuzz campaign exists to break the attach path, and it has found real bugs — six in roughly thirty cells on one run, all in the attach path. Those are fixed. It would be surprising if there were none left.

What happens if I lose the terminal mid-attach?

The probe carries on and then expires on its own, because the expiry and the rate breaker are counters inside the injected library rather than policy in the tool. Killing the CLI, losing the network or crashing the collector does not extend a probe's life.

If injection itself is interrupted, threads that land in the patching window hit a SIGTRAP safety net and are rewound to re-execute the completed jump.

Does it work on a stripped binary with no source?

Yes. The nginx on the Proof page is a stock distro package where nm reports no symbols at all — the probe sites came from exported dynamic symbols. Where a separate debug file exists it is used, including via .gnu_debuglink, and where DWARF exists you additionally get line-precision sites and capture by variable name.

What does it cost when nothing is probed?

Nothing. There is no DBT engine, no code cache and no dispatcher in the process — unprobed paths execute the original instructions. The cost is a 5-byte jump at the sites you chose and the work your capture asks for.

Do I need root? Does it work in a container?

You need to be able to ptrace the target, which normally means the same user or CAP_SYS_PTRACE. One thing that costs people time: the target reads the probe library and the spec as its own user, so both must be readable by it — a probe under /root fails against an nginx worker running as www-data, which is exactly how the Proof page's first attempt failed.

Can I remove it cleanly, and prove I did?

--remove-all restores the original bytes and the process keeps running. On the Proof page that is 36 of 36 nginx workers still alive and serving HTTP 200 afterwards, on the same pids.

Where does the captured data go?

A shared-memory ring in the target, drained to NDJSON on the same machine. Nothing is transmitted anywhere. There is no server component in the attach path at all — which is also why the guardrails had to be implemented inside the target rather than in a control plane.

What is it worst at?

Capturing secrets, because redaction and blocklists do not exist yet — a probe on an auth path will capture tokens. Recording a program that crashes, which does not work today. Anything needing conditional capture, .NET, Windows, macOS or serverless. Every one of those is stated on this site rather than left to be discovered.

Also in the box

Beyond attaching

Reverse debugging

Record a run to a .fzr file and step backwards through it in GDB with reverse-continue and reverse-step. Works on small recordings; the indexer currently caps at 4,096 events.

19 analysis passes

Cache simulation, coverage, flamegraphs, memory tracing, heap profiling, syscall mix, branch profiling — all over one unmodified recording.

5 architectures

x86-64, aarch64, riscv64, arm32 and mips64. Graviton is first-class, not an afterthought.

For TAC, field engineering and support

The box is the
customer's, and it
cannot be restarted

A support escalation is a debugging problem with every useful option removed. You did not build this deployment, you cannot reproduce it, the binary is a release from months ago, and the one action that would tell you the most — restart it with more logging — is the one action that destroys the evidence and burns the maintenance window.

This is the case flux-dbg was shaped by, and it is a harder case than the one developer-observability products are built for.

Why the usual tools don't apply here

Every assumption a debugger makes is false on a customer box

What a debugging tool usually assumesOn an escalation
An agent was loaded when the process startednobody knew there would be an incident
You can redeploy with more loggingit's a change-controlled appliance
You can reproduce it locallyit needs their traffic, their config, their peers
The binary has symbols and a source treeit's a stripped release build
You can attach GDB and stop the processstopping the daemon is the outage
The customer will run whatever you sendthey will ask exactly what it does first

flux-dbg holds under all six. It attaches to an already-running process, it resolves probe sites from a stripped binary's own symbols or DWARF, it never stops the process, and the patch is removed with one command. What you hand the customer is a spec file and a command, not a debug build.

The same escalation, twice

What the week looks like either way

Without

Four round trips, minimum

Day 1 — customer reports it. You ask for logs. The logs don't have the field.

Day 3 — you build a debug binary with extra logging. Change control needs a window.

Day 8 — the window happens, the daemon restarts, and the condition doesn't reproduce for another week.

Day 15 — it reproduces. The new log line tells you that you instrumented the wrong function.

With

One session, no restart

Hour 1 — compile a spec against the exact build they're running. The build-ID makes it impossible to apply to the wrong image.

Hour 1 — they run one command with a visible expiry and rate limit. The daemon keeps serving.

Hour 2 — real arguments under real traffic. Wrong guess? The probe expires and you place another one.

Hour 3 — evidence bundle with the build-ID and the spec that produced it, so the fix can be argued from data.

The difference isn't speed, it's the cost of being wrong. When a hypothesis costs a maintenance window, you only test the ones you're confident about — which is the opposite of how debugging works.

The workflow

From escalation to evidence bundle

01 · SCOPE

Compile a spec against the exact build the customer is running. The build-ID is baked in, so it cannot be applied to the wrong image by accident.

02 · BOUND

Set the blast radius up front — --ttl-sec and --max-hz. The customer can read both off the command before they run it.

03 · COLLECT

Probes fire under their real traffic. Records land as NDJSON: arguments, return values, struct fields, timings.

04 · LEAVE

The probe expires on its own even if nobody comes back, and --remove-all restores the original bytes.

Step 02 is the one that makes step 03 possible. "Attach a debugger to my production router" is a conversation that ends quickly unless the answer to "and what if you're wrong?" is a number the customer can see.

Where the AI part comes in

First-line triage that produces evidence, not a guess

Most escalation time is spent before anyone learns anything: narrowing which subsystem, deciding what to ask the customer for, waiting a cycle for the answer. An agent with 44 MCP tools and a live process can do that narrowing against the system itself.

It states a hypothesis first

The catalog is built from the target's own DWARF and safety-filtered, so the agent selects from probeable functions rather than inventing addresses. What it installs is auditable before it runs.

It works inside limits it cannot widen

Both attach tools default to a 15-minute expiry and a 20,000/sec per-site ceiling. If a policy on the box is tighter, the tighter one wins — an agent cannot raise a ceiling an operator set.

It hands back evidence, not prose

Captured arguments and timings with the build-ID and the spec that produced them. A human engineer can check the claim rather than trusting the summary.

It stops being expensive to be wrong

A wrong hypothesis costs one expired probe, not a maintenance window. That is what makes iterating on a customer's live system reasonable at all.

Fit

Where this lands hardest

Network operating systems and routing daemonsC and C++, long-lived, restart = outage
Appliances and on-prem deploymentsno redeploy path, change-controlled
Databases and storage enginesrestart costs a cache warm-up, or worse
Embedded and edge fleetsaarch64 is first-class, not an afterthought

Notably, these are the deployments where the software is compiled to native code — the category the managed-runtime products do not cover at all. Comparison →

Use cases

What people actually
reach for this to do

Organised the way the decision gets made — by the outcome someone is trying to buy, and by whose week it changes. Each one says what it needs from you, because the honest answer is not always "nothing".

By outcome

The thing you are trying to make true

Close a bug that only happens in production

Attach to the process that is misbehaving, capture the arguments it is really being called with, detach. The state survives, because nothing restarted.

Needs: the ability to ptrace the target, and its build available to compile a spec against.

Cut log volume you added defensively

Log lines added "just in case" are a permanent monthly cost paid for a failure that may never recur. Instrumentation you can add at the moment you need it is the alternative to predicting in advance.

Needs: a willingness to attach at incident time — the cost only comes back if the answer is "no, add it to the build".

Debug a language your APM doesn't cover

C, C++, Rust, Go, Erlang, PHP, Ruby — plus the JVM and Node families. One attach model across all fifteen, rather than a different mechanism per runtime.

Needs: Linux, and one of those fifteen. Not .NET.

Answer "what is it doing right now" without a deploy

Span-kind probes fold into a per-function digest — count, p50, p95, max, with sample arguments — so "which call is slow" is a command rather than an afternoon.

Needs: load flowing through the path you probed. It observes; it does not synthesise traffic.

Instrument a hardened or AOT-compiled binary

Targets with Intel CET shadow stacks armed, and GraalVM native-image executables with no JVM in them. Both verified rather than argued.

Needs: nothing special — the CET path is selected at runtime when a shadow stack is detected.

Give an AI agent something real to look at

44 MCP tools, so an agent can place a probe where its hypothesis says the answer is and revise from what actually happened — inside limits it cannot widen.

Needs: an MCP-capable client, and a decision about what you are comfortable letting it attach to.

By role

Whose week this changes

TAC & field engineering

The strongest fit, and not a coincidence

Customer box, shipped build, no reproduction, no restart permitted. Every assumption a debugging tool normally makes is false, and this is the case the tool was shaped by.

The TAC workflow →
SRE & on-call

Add the field the logs don't have, now

Mid-incident instrumentation on a running service, with an expiry and a per-site rate breaker enforced inside the target so the probe cannot become the second incident.

Backend & systems engineers

Stop reproducing, start observing

The bug that needs their traffic, their config and their timing is the one you cannot rebuild locally. Probe it where it lives instead.

Platform & infrastructure

One mechanism across a polyglot fleet

Fifteen languages, five architectures, aarch64 first-class. One tool to approve, one set of guardrails to reason about, rather than a separate agent per runtime.

Missing from this list, deliberately: engineering leaders. Every honest thing to say to them is about fleet rollout, RBAC, audit and redaction — none of which exists yet. What does →

Pricing

Free until it's
someone's job

An engineer debugging their own systems should never see a bill or a quota. The line is drawn at fleet scale, where this becomes infrastructure that an organisation depends on and expects someone to answer the phone about.

Free

$0

Everything the tool does, on up to a small number of hosts. Not a trial, not a feature-gated demo, and not time-limited.

· Every language, every capability
· Live attach, record/replay, all analysis passes
· Full MCP surface and the agent loop
· Production guardrails
· Community support

Limit: a fixed host count, stated plainly rather than discovered when you hit it.

Enterprise

Licensed

For fleets, and for the things an organisation needs that an individual does not.

· Unlimited hosts
· Redaction, blocklists and audit on the roadmap
· RBAC and SSO on the roadmap
· Air-gapped and offline licensing
· Support with a response time attached to it

Gate: a license file, checked locally.

The roadmap markers are deliberate. The enterprise controls that gate real deals — redaction, blocklists, audit, RBAC — are not built yet, and are listed here as what the license will cover rather than as what it covers today. What exists →

How the limit is enforced

A registration check, and what it does with your data

The free tier's host limit is enforced by a lightweight registration check at attach time. Enterprise builds then validate a license file for entitlements beyond the free ceiling. Being specific about this matters more than the mechanism does, so:

What is sentan install identifier, version, and a host count
What is never sentsource, symbols, captured values, probe specs, hostnames
Where captured data goesnowhere — it stays on your machine, in your ring buffer and your files
Enterprise licensesvalidated locally against the license file; air-gapped supported

The deployments this tool is best at — appliances, network operating systems, on-prem fleets — are frequently offline by policy, so an offline path is a requirement of the design rather than an exception to it.

The economics, without a fabricated ROI figure

What the alternative actually costs

Competing products put a modelled return on this page — a percentage and a net present value over three years. There is no such study here, and inventing one would undermine everything else on this site. What can be stated is the shape of the cost being avoided, which you can price for your own organisation better than anyone else can.

A restart to add logginga change window, a cache warm-up, and the loss of the state you were investigating
A reproduction attempt that failsdays of calendar time, and the ticket ages while nothing is learned
Log lines added "just in case"permanent volume, paid monthly, for a failure that may never recur
An escalation to the engineer who wrote itthe most expensive person in the company, context-switched

The free tier costs nothing, so the only real question is whether it works on your binaries. That is answerable in an afternoon, which is deliberately a shorter path than a procurement conversation.

Not yet available

Neither tier is purchasable today. The free tier is the whole tool, from GitHub, right now. The enterprise tier is a statement of where the line will be drawn, published early so it is not a surprise later.

Get the free tier See the TAC workflow

Coverage

Fifteen languages.
One attach.

Not fifteen integrations, fifteen agents, or fifteen sets of documentation — one mechanism, because it instruments the binary rather than the runtime. Every language below has a live-attach test in the repository, and the ones that don't work are listed beside the ones that do.

Compiled to native

The category the agent-model products cannot reach at all, because there is no bytecode to rewrite.

LanguageHow probe sites are foundStatus
CDWARF, or exported symbols on a stripped binaryverified
C++DWARF with demangling; inline sites and optimiser clones handledverified
RustDWARF including namespaced paths and generic instancesverified
GoDWARF; Go's own linker layout handled explicitlyverified

Managed runtimes

Where the agent-model products live. flux-dbg reaches these too, by the runtime's own instrumentation interface rather than by binary patching.

LanguageMechanismStatus
JavaJVMTI agent, Byte Buddy / ASM weaving, live add and removeverified
KotlinJVM bytecodeverified
ScalaJVM bytecodeverified
GroovyJVM bytecodeverified
ClojureJVM bytecodeverified
Node.jsV8 Inspector Protocol; breakpoint-based capture, CJS and ESMverified
BunJavaScriptCore Web Inspector; breakpoint capture (no LiveEdit — JSC has none)verified
TypeScriptVia Node.js, including bundled and source-mapped buildsverified
PythonNative probes on CPython plus a Python-level pathverified
RubyNative probes on the interpreterverified
PHPNative probes on the interpreterverified
ErlangBEAMverified

The interesting edge

GraalVM native-image

Bytecode instrumenters have a hard edge: ahead-of-time compilation. GraalVM native-image compiles Java to a native ELF with no JVM and no bytecode at runtime, and every mechanism a JVMTI-based agent owns stops existing. Lightrun's own documentation says so plainly — "At this time Lightrun doesn't support native compilation with GraalVM", and their known-issues list adds that GraalVM works "only when used as a JVM without native images".

A native-image executable is an ELF with DWARF, which is what flux-dbg already instruments for C, C++, Rust and Go. The managed/native distinction isn't load-bearing here — and this is now tested rather than argued.

GraalVM used as a JVM (no native image)verified — it's just the JVM path
GraalVM native-image executablesverified — probed live, arguments correct

A service compiled with native-image, running, with no JVM in its link map. The probe resolves the Java method by symbol name out of the binary's own DWARF, patches it live, and captures its arguments:

394 captures, every argument pair satisfying the original Java arithmetic. Verbatim. This is a Java program being debugged live with no JVM present at all — the case a bytecode instrumenter cannot reach by construction.

And by source line, and by Java variable name

The symbol-name path above is the blunt one. native-image also emits full DWARF, so a probe site can be chosen by source line and its captures named in Java rather than in registers:

Getting here needed a real fix. native-image emits each Java method as a declaration/definition split — the DIE holding the address carries no name at all, only a reference to the declaration that does — so name lookup had to learn to follow DW_AT_specification and DW_AT_abstract_origin. Not a GraalVM quirk: an out-of-line C++ member definition has the identical shape, and the regression test uses that so it runs without GraalVM installed.

One practical note: GraalVM names parameters __0, __1 unless the class was compiled with javac -parameters, with which the real Java names survive into the native image.

Not supported

.NET / C#no — the one runtime the agent-model products cover and this doesn't
Windows and macOS targetsno — Linux only
Serverless (Lambda, Cloud Functions)no — the attach model assumes a pid you can reach

Evidence

Run against software
nobody here wrote.

There are no customer logos on this page, because there are no customers yet. There is something a logo can't give you: a list of real third-party programs this has been attached to, what was captured, and what broke.

Captured today, verbatim

A stock nginx that nobody prepared

Distro package, stripped — nm reports no symbols. Running as www-data, up for a day and a half, doing nothing special. Probe sites came from exported dynamic symbols alone.

60 requests in, 60 captures out, on a binary with no symbol table.

The part that wasn't planned

It refused to patch a process it wasn't sure about

The first attempt matched every process whose command line said nginx: worker. That swept in workers from a container image — a different nginx build at different offsets. A spec's probe sites are file offsets valid only for the build they were compiled against, so patching those would have written a jump into the middle of an unrelated instruction, in a live process, as root.

Nobody staged that. It was an ordinary mistake made while recording the transcript above, and the guard caught it. That is the entire argument for this class of tool being conservative by default: the operator will be wrong sometimes, usually at three in the morning.

Third-party software attached to

What was captured, per target

TargetLanguageResult
nginxC36/36 workers · 60/60 requests captured, stripped binary
NumPyC / Python1,018,508 / 1,018,508 hits — exact 1:1 match
JVM workloadJava97.8M captures, 0 drops
CeleryPython125K hits · 200/200 tasks
PyTorchC++ / Python131K hits, both sides of the boundary
etcdGo20/20 hits
CockroachDBGo18 hits on connExecutor.execStmt from 5 SQL queries
containerdGo5/5 hits — after fixing a real probe-side stall
TiKVRustlive-verified after fixing Rust namespace resolution
DatabendRustattaches cleanly — two real injection bugs fixed to get there
LangChain.jsNode.js240 hits, correct args, correct async
FastifyNode.js11/11 hits, CJS and ESM
FRR (bgpd)Clive BGP session between namespaces, probed
PostgreSQL, Redis, HAProxy, Envoy, MinIO, Prometheus, Grafanamixeddemo suites in the repository

Several of those rows exist because the attach failed first and a real bug was fixed. Databend surfaced a wrong-libc-mapping bug and a crash on fprintf to a null stderr; containerd surfaced an unsafe fopen inside the injected library that stalled it. Attaching to software you didn't write is how those get found.

Measured, not "negligible"

Overhead, with its methods attached

Overhead depends entirely on the path, so a single headline number would be dishonest. These are the ones that exist, each with the caveat that belongs to it.

Unprobed code paths, probe attachednative — no DBT, no dispatcher
Java probes, JVMTI path~0.5–1.6%
Node.js breakpoint-fallback path~11–15% — a lower bound, not a figure
Full record/replay, compute workload~3×
Full record/replay, syscall-heavy workloadup to ~40×

What record/replay does not do

The live-attach path is the mature one. Record/replay is real and useful, and it has limits that are easier to find than the marketing version of this page would suggest — so here they are, with the permanent ones marked as permanent rather than dressed up as roadmap.

Replay a recording, deterministicallyworks
Reverse execution in GDB on a small recordingworks
19 analysis passes over one recordingworks
Recordings above 4,096 eventsindexer caps there — a trivial program can exceed it
Recording a program that crashesnot supported, and not planned — a limit of the recorder's design, not a missing feature
Short-lived tools (ls, cat)best-effort; the tool says so on load
Record/replay under Intel CETnot yet — see Intel CET

Concretely, on a nine-line C program that dereferences a corrupted pointer: the recorder printed the program's own output nineteen times instead of once, and the resulting recording held over two hundred thousand events against a four-thousand-event index cap. The event cap is an ordinary bug. The crash behaviour is not — it follows from how the recorder executes, and it is not going to be fixed, so "record the crash and walk backwards to the write that caused it" is not a thing this tool will do. An earlier version of this page implied otherwise, which is the kind of claim the rest of the site exists to avoid.

The Node figure is deliberately unflattering: it is the fallback capture path, measured against a target whose non-deterministic relink diluted the result, so it is published as a lower bound. The competing products publish no quantified overhead number at all — across 351 pages of their documentation, every claim is "minimal", "negligible" or "zero overhead".

Production guardrails

A probe should never
become the incident

Two failure modes make live instrumentation frightening: the probe someone installed and forgot, and the probe placed on a path that turns out to run per-packet. Both are handled inside the target, and both can be set on a process that is already running.

Enforced in the target, not in the tool

This distinction is the whole design. A limit held by the CLI, the server or the agent is a limit that disappears the moment any of them does. These are counters inside the injected library — they survive the network, the collector and the operator.

The baseline row exists so the others mean something: an identical probe with no limit keeps climbing across the same windows.

The three rules

Expiry

--ttl-sec N — probes go inert N seconds after they load. The patch stays until you detach; it simply stops costing anything. For the engineer who attached to answer one question and got pulled into a meeting.

Rate breaker

--max-hz N — any single site over N hits/sec is disabled and stays disabled. Per site, so one runaway probe doesn't take the others with it. Sampling already caps volume, but only if someone predicted the rate.

Limits only tighten

When the service owner has set a ceiling and an attach asks for more, the tighter value wins. An attach cannot widen a policy — which matters most when the thing attaching is an autonomous agent.

Refusals are a feature

Spec build-ID doesn't match the running binaryrefuses to patch
A guardrail was requested but can't be writtenrefuses to attach, and says why
A capture faults at runtimeemits null, target survives
Probe site outside the mapped modulebounds-checked, rejected

A refusal that explains itself costs an engineer thirty seconds. A hopeful patch into the wrong build costs them the process.

Missing, and honestly so

  • No PII redaction. A capture emits whatever is at the address. Designed, not built.
  • No blocklists. Nothing yet prevents instrumenting a sensitive package or class.
  • No audit trail. Who installed which probe, when, against which build-ID — not recorded.
  • No RBAC. Attaching is direct: whoever can ptrace the process can probe it.

These are the current priority, in that order.

Hardware control-flow integrity

The CPU thinks
you're an attacker.

A shadow stack is a second copy of every return address, maintained by hardware, that software is not allowed to write. It exists to kill return-oriented programming — and from the processor's point of view, a naive trampoline is indistinguishable from an ROP chain. So it kills that too.

This is the failure no bytecode instrumenter ever meets, and every binary patcher eventually does.

This is the failure mode no bytecode instrumenter ever meets and every binary patcher eventually does. fr trace handles it, on hardware that actually enforces it.

Why it breaks trampolines

The transfer that looks like an attack

When a probe relocates a CALL out of the patched function, the obvious implementation synthesises the call so the callee returns straight back to native code and never sees the trampoline. That's deliberate — it keeps the probe transparent to a frame-walker or Go's garbage collector. It looked like this:

push rax
movabs rax, native_return
xchg  rax, [rsp]
jmp   qword [rip+0]        ← and the indirect forms end in `ret`

The push touches the normal stack only. The ret then consumes a shadow-stack entry belonging to someone else. That is precisely the shape a shadow stack exists to reject, so the hardware faults — SEGV_CPERR, and not at the probe: at the return of the function the probed one called.

The illusion this created

It only bites a function that makes calls. Basic-block analysis treats CALL as a terminator, so a leaf function's terminator is a RET and nothing gets synthesised. The multithreaded reproduction happened to probe a non-leaf; the single-threaded one a leaf. That produced a convincing "only fails with threads" story that cost real time. Threads were never the variable.

The fix

Let the hardware do it

When the shadow stack is armed, emit the real instruction and jump back afterwards, so the CPU pushes both stacks and the callee's ret matches:

<relocated call>           ← hardware pushes BOTH stacks
jmp   qword [rip+0]
.quad native_return

All three forms: direct via call qword [rip], register-indirect verbatim, memory-indirect with its RIP-relative displacement rebased. The cost is transparency — the callee now sees a trampoline return address — so it is applied only when CET is actually armed, checked at runtime. Every other target keeps the old, fully transparent behaviour.

Verified where it counts

On silicon that enforces it

This could not be verified on an ordinary cloud instance. On a virtualized Sapphire Rapids host, arch_prctl(ARCH_SHSTK_ENABLE) returns EOPNOTSUPP — the silicon has CET and the kernel supports it, but the hypervisor withholds it. That's the third feature blocked this way, after the PMU and Intel PT. Verification ran on bare metal.

The right-hand column is after the fix. A real routing daemon, with shadow stacks armed on four threads, holding a live BGP session across being probed.

Two more things CET taught us

Findings worth having anyway

Linux has no user-mode IBT

The kernel's user CET interface exposes only ARCH_SHSTK_SHSTK and ARCH_SHSTK_WRSS — there is no IBT bit. So on a stock distro today, indirect-branch tracking has no practical effect on user code, whatever the binary's notes say.

One unmarked library disables it everywhere

glibc arms the shadow stack only if the executable and every loaded DSO is marked. A single unmarked dependency silently turns CET off for the whole process — so "we compiled with -fcf-protection" and "CET is on" are different statements.

It refuses rather than guesses

Before the fix existed, the compiler happily built an exit probe against a CET binary and reported "0 warnings, 0 skipped". Now the mechanism is chosen from the binary's .note.gnu.property automatically, and an attach that can't be made safe is refused.

Emulators are not hardware

Intel SDE was a faithful stand-in for the fault, but it arms from the binary's notes rather than the glibc request, and implements the old ARCH_CET_* interface rather than the ARCH_SHSTK_* one real kernels use. A fix validated on SDE is not a fix validated on Linux.

What is still broken under CET

fr trace — live probe attach, entry / exit / spanworks, hardware-verified
fr record — full record/replay under DBTstill fails on a CET-enforcing target
aarch64 GCS — the same question on ARMnot investigated

The recorder's translator emulates CALL as push; jmp and RET as pop; jmp, so the data stack moves and the shadow stack never does. Two candidate fixes are written up and neither is finished. If you need record/replay on a shadow-stack-enforcing binary today, this is not ready — the live-attach path is.

Agent-driven debugging

Let your agent
watch it happen.

An AI agent reading logs is reasoning about the past, in text someone decided to write down in advance. Give it the running process instead: it places a probe where its hypothesis says the answer is, watches what the code actually does, and revises — the same loop a good engineer runs, at machine speed.

Managed-runtime agents can do this for four languages. This does it for fifteen, native included — because the agent drives the same attach a human would.

Not a diagram — a transcript

The loop, exactly as it ran

Below is a real run of the reference agent against a live PostgreSQL backend, over the Model Context Protocol. It was handed a symptom and a pid, nothing else. Every line is the agent's own — its stated reasoning, the tool it chose, and what came back.

The agent picked the catalog, read the latency-tagged entries, installed a timing probe, then installed WaitOnLock specifically to confirm-or-kill the lock-contention hypothesis. That sequence is a decision, not a script.

What it can reach for

44 tools, one live process

Build a probe catalog

From the target's own DWARF, safety-filtered — so the agent selects from functions known to be probeable rather than inventing an address.

Install on a live pid

Attach a probe to a running process, always inside a default expiry and per-site rate ceiling it cannot widen.

Read what came back

Drain the ring, or fold spans into a per-function latency digest — count, p50, p95, max, with sample arguments.

Time-travel a recording

reverse-continue, watchpoints and backtraces over a recorded run, within the recording sizes it currently indexes.

The part that took the work

An agent's limits are enforced by the process

The loop is the easy half. The hard half is what happens when the caller is autonomous and points it at production. Both attach tools default to a 15-minute expiry and a 20,000/sec per-site ceiling — limits, not “unlimited” — and they are counters inside the target, so they hold even if the agent's context is truncated mid-investigation.

Competing agent skills spend a section instructing the model to clean up the actions it created. This does not need one: a limit enforced by the process is not the same as hygiene requested of a model. The guardrail model →

Where it has really run

And when it captures, this is what it sees

The transcript above is the loop. This is the evidence end of it, from a real attach to SONiC's control plane — the kind of native, stripped, containerised target the agent-model products cannot reach at all. One operator shuts a port; the capture shows every object the control plane touched in response:

See the full SONiC run Get the skills package

Honest about this one too

What the agent path does not do yet

  • It has no redaction. An agent placing probes on a live system captures whatever is at the address, including secrets — the same gap the rest of the tool has, and it matters more when the caller is autonomous.
  • Its conclusion is only as good as its captures. When a probe returns no data, a weak agent will conclude from the absence. The reference loop is real; treat any single run's verdict as a hypothesis, not a verdict.
  • The packaged skills ship without telemetry. That is a choice, and the opposite of the norm — but it means adoption is invisible to us by design.

Comparison

Including the rows
where they win.

Two different tools get compared to this one, for opposite reasons. Lightrun, Rookout and Datadog Dynamic Instrumentation debug production — but only a process their agent was loaded into first. Tracy is the profiler serious C++ teams reach for — but only code it was compiled into. Both are good, and a table that only flattered us would tell you nothing, so each of these shows the rows they win.

Agent-model productsflux-dbg
Attach to an already-running processno — agent must be loaded at startyes
Languages4 managed runtimes15, native and managed
C, C++, Rust, Gonoyes
.NET / C#yesno
GraalVM native-imageno — documented limitationyes, verified
Metrics primitiveJava onlyall languages
Conditional captureyesnot yet
PII redactionyesnot yet
Blocklistsyesnot yet
RBAC, SSO, SCIM, audityesno
Published overhead numbersnone quantifiedper path, with method
Targets with Intel CET shadow stack armedn/a — bytecode instrumenters never touch a shadow stackyes, hardware-verified
Record & replay deterministicallynoyes, with size limits
Serverless (Lambda, Cloud Functions)yesno
Observability integrations16+ documentedOTLP export
Certifications (SOC 2, ISO 27001)yesno

The other foil — for the native story

Versus Tracy, the profiler you compile in

Tracy is the tool C, C++ and game developers actually reach for, and it is genuinely excellent at what it does. It is also the cleaner comparison for our native story than any managed-runtime debugger — because it shares our machinery (a lock-free ring, rdtsc timestamps, perf_event_open sampling) and diverges on exactly one axis that decides everything: Tracy has to be compiled into your source; we attach to the binary after the fact.

Tracyflux-dbg
How instrumentation gets insource macros + recompile + relinkattach to an unmodified binary
A process already runningno — must have been built with Tracyyes
Needs the sourceyesno — stripped binaries work
Marking what to tracehand-placed ZoneScoped in sourceany function, by symbol or address
Real-time visual profiler UIyes — best in classno — NDJSON + a source viewer
Lock-free ring, ns timestampsyesyes
Kernel sampling (perf_event_open)yesyes
GPU profiling (Vulkan, D3D, CUDA…)yesno
Built foryour app, in developmentsomeone else's service, in production

This is not a tool we are trying to beat. It is where the line sits: Tracy is what you want while you are building the thing; this is what you want once it is running somewhere you cannot rebuild it. Their real-time UI is the bar for any visualization we add past NDJSON — and we say so plainly rather than pretend the gap is not there.

Where the comparison isn't apples to apples

Hit ceiling vs rate breaker

Their limit is a total — a default of 50 hits and then it stops. Ours is a rate. Theirs suits snapshot debugging; ours suits continuous tracing. Neither is a better version of the other.

Server-mediated vs direct

Every action of theirs goes IDE → management server → agent, which is where their RBAC, audit and redaction live. We attach directly: lower friction for one engineer, unbuyable for an enterprise. Same fact, two readings.

Prepared-in-advance vs attach-after

Every one of these had to be set up before the incident — an agent preloaded (Lightrun, Datadog), or the source recompiled (Tracy). We attach to what is already running. That is one wall, reached three ways, and it is the whole reason this tool exists.

"Debug live production" means two things

For them: add an action to a process already carrying their agent. For us: attach to a process that knows nothing about us. Both are true uses of the phrase.

Comparisons cite public vendor documentation captured 2026-09-10 and describe those products as their own docs describe them.

Blog

Notes from
building a debugger
that attaches.

Every post here is a real bug, a real fix, or a real mistake — with the numbers and the code, including the parts that did not flatter us. That is the same rule the rest of this site runs on.

← all posts Intel CET

The CPU thought we were an attacker

hardware control-flow integrity

A shadow stack is a second copy of every return address, kept by the hardware, that software is not allowed to write. It exists to make return-oriented programming impossible. It also, it turns out, makes a naive trampoline injector impossible — for exactly the same reason.

When a probe relocates a CALL out of the function it is patching, the obvious implementation synthesises the call so the callee returns straight back to native code and never sees the trampoline. We did that on purpose, to stay invisible to a stack walker or Go's garbage collector. It looked like this:

push rax
movabs rax, native_return
xchg  rax, [rsp]
jmp   qword [rip+0]      ← the indirect forms end in `ret`

The push touches the normal stack only. The ret then consumes a shadow-stack entry that belongs to someone else. From the processor's point of view that is indistinguishable from an exploit, so it does what it is designed to do: SEGV_CPERR, and not at the probe — at the return of the function the probed one called.

It only bites a function that makes calls. That one fact disguised the whole bug.

Here is the part that cost real time. Basic-block analysis treats CALL as a terminator, so a leaf function's terminator is a RET and nothing gets synthesised. Our multithreaded reproduction happened to probe a non-leaf function; the single-threaded one happened to probe a leaf. That produced a completely convincing story — "it only crashes with threads" — and we chased concurrency for a while. Threads were never the variable. Leaf-versus-non-leaf was.

The fix

When the shadow stack is armed, emit the real instruction and jump back afterwards, so the hardware pushes both stacks and the callee's ret matches. It costs the transparency we were protecting — the callee now sees a trampoline return address — so it is applied only when CET is actually armed, checked at runtime. Every other target keeps the fully transparent path.

Verified on real silicon that enforces it — a bare-metal Sapphire Rapids box, because a virtualized instance withholds CET entirely:

mt_shstk, 4 threads    entry   4 records DEAD → 97,541 records ALIVE
real FRR bgpd, shadow stacks armed on 4 threads, BGP session up:
                       1 record then DEAD → 3 records, correct, session still up

This is the failure mode no bytecode instrumenter ever meets, and every binary patcher eventually does. It is also the clearest example of why we run on hardware that enforces the thing, rather than an emulator that approximates it — an emulator arms from the binary's notes, not the running CPU, and would have told us a comforting lie.

← all posts
← all posts GraalVM & a silent bug

394 correct, 0 wrong — after the attach found nothing

native-image · DWARF

GraalVM's native-image compiles Java ahead of time to a native ELF with no JVM and no bytecode at runtime. Every mechanism a bytecode instrumenter owns stops existing. Which is exactly why it should be easy for a tool that instruments the binary — and it was, after it wasn't.

The first attach captured nothing. Not an error — worse, a success that produced an empty trace, which reads as "this function never ran." The function ran fine. The tool just could not find it.

native-image emits a stripped executable plus a .gnu_debuglink pointing at a debug file beside it. Our symbol resolver had a docstring claiming, since the day it was written, that it "honours the .gnu_debuglink convention." It did not. It only ever looked under /usr/lib/debug. The documented behaviour was simply never implemented.

It fails silently, which is how it survived. Zero probes emitted, attach reports success, trace reads as "never ran."

And it was not a GraalVM problem at all. objcopy --only-keep-debug plus --add-gnu-debuglink is the standard way to split debug info — so every locally-built split-debug binary had been hitting this, quietly, the whole time. GraalVM just happened to be the target that made us look.

The check that would have caught it

Once .gnu_debuglink resolution worked, the Java method resolved by name and the probe fired. The proof is not the record count, it is the arithmetic. The source is handleRequest(i, (i % 7) + 100), so every captured pair must satisfy size == id%7+100:

$ fr trace collect $PID
  {"func":"handleRequest","vars":{"id":227,"size":103}}
  {"func":"handleRequest","vars":{"id":228,"size":104}}
  argument check: 394 correct, 0 wrong

394 out of 394. That rules out the failure where you patch something and get plausible-looking numbers. The regression test now asserts the thing that was actually broken — that the compiler emits one probe, not zero — and it uses an ordinary C++ out-of-line member definition, which produces the identical debug-link shape, so it runs anywhere without GraalVM installed.

← all posts
← all posts An incident, ours

How I took down my own server proving a fix worked

postmortem · the honest kind

This one is on us, and it is here because a site that only publishes its wins is not one an engineer should trust. Two failures, one after the other. The first was a real bug. The second was worse, and it was a judgement error.

The runaway

Roughly a thousand processes from our own test suite, each parented by the last, had been alive for nearly a full day, holding a 36-core box at a load average of over a thousand. Nobody's monitoring caught it — a human looked at the machine and asked.

The test binary was innocent: one fork(), parent accepts, child connects, parent waits. But the suite records and replays it, and replaying a program that forks re-executes the fork — so it chained instead of running once. Killing the thousand processes went cleanly.

The part I got wrong

To prove the containment guard I had written actually worked, I ran a deliberate fork bomb, as root, on the same box that had just recovered. It exhausted the PID table. sshd stayed up and port 22 kept accepting connections, but no session could start, because nothing could fork(). It needed a console reset.

A guard validated in the wrong privilege context is not validated. It only looked like it worked.

The guard's first version used ulimit -u. But RLIMIT_NPROC is bypassed by any process holding CAP_SYS_RESOURCE — which is to say, by root, which is exactly how those tests run. So the cap protected nothing in the one case that mattered, and I proved that by running the bomb as root. The corrected guard uses cgroup v2 pids.max, which is enforced for root.

What we kept

Two rules, written down so they don't recur. Never validate a destructive failure mode on infrastructure you cannot power-cycle — a throwaway VM exists for that. And a symptom worth recognising fast: ping fine, port 22 accepting, SSH hanging at the banner is PID-table exhaustion, not a dead host — the disk and everything on it are intact; it needs a reset, not a rebuild.

The tool's own recovery worked perfectly. The operator running it made an avoidable call. Both halves are true, and both are in the record.

← all posts
← all posts A new language

A 16th language in an afternoon, because we'd already met the hard part

JavaScriptCore

Bun looks like Node from the outside and is nothing like it underneath. It embeds JavaScriptCore, not V8. So the entire Node.js source-level tracing path — the Inspector connection, the source-rewrite trick — simply did not apply. It still shipped in an afternoon, and the reason is a small lesson about paths not taken.

The rule on this project is: don't add a language to the list on a hope. So before writing a line of the backend, the question was whether JSC's inspector even exposes what we need. It exposes almost all of it — breakpoints by source location, evaluate-on-frame, script parsing — and is missing exactly one primitive: setScriptSource, V8's live-edit.

The one primitive JSC lacks is the one we'd already abandoned on V8, for its own reasons, two months earlier.

V8's live-edit rejects ES modules and has a subtle by-value-capture problem, so the Node path had already moved to a breakpoint-based capture — set a breakpoint at the source line, read the arguments off the paused frame, resume. That is precisely, and only, what JavaScriptCore offers. The dead primitive was the one we didn't use; the live one was the one we depended on.

Two engines, two gotchas

The backend that resulted is 167 lines, against 1,510 for the Node one — JSC needs none of V8's source-map machinery. Getting it to fire took two fixes, both echoing older Node lessons:

  • An exact script URL arms a JSC breakpoint. A regex-matched URL reports a resolved location and then never fires — a resolved-but-dead breakpoint, the most confusing kind.
  • JSC does not populate hitBreakpoints on a pause the way V8 does, so the probe has to be matched by the paused location — and the breakpoint must sit on an executable body line, not the function declaration line, which does not run per call.

Then it captured — arguments and a closure local, by name, from a live Bun process, every value matching the source arithmetic. Sixteen languages, and the newest one earned its place by a passing test, not an announcement.

← all posts

Get started

Attach to something
you didn't write

The interesting first test isn't the example binary. It's a service you don't have the source tree open for, that's already running, that you'd rather not restart.

Build it

Dependencies are GCC, zlib and pthreads. Linux only.

$ git clone https://github.com/rajeshgangam/flux-dbg
$ cd flux-dbg && make -j$(nproc)

Probe a running process

# describe what to capture
$ cat > req.tpc <<'EOF'
{"service":"nginx","name":"req","probes":[
  {"kind":"entry","func":"ngx_http_process_request","id":"req",
   "captures":[{"name":"r","reg":"arg0"}]}]}
EOF

# compile it against the binary that pid is running
$ python3 tools/tpc-compile req.tpc /usr/sbin/nginx -o req.bin

# attach, with a blast radius
$ fr trace attach $PID --spec req.bin --ttl-sec 300 --max-hz 20000

# read what came back, then put it back the way you found it
$ fr trace collect $PID
$ fr trace attach $PID --remove-all

Two things that cost real debugging time if you skip them: the target reads the probe library and the spec as its own user, so both need to be readable by it; and the spec must be compiled against the exact build the process is running, which the build-ID check enforces.

Or record and go backwards

$ fr record --recording-file crash.fzr -- ./my_program arg1
$ fr replay crash.fzr        # GDB, with reverse-continue

What to expect

It's a research-grade tool

Single engineer, no company behind it, no support contract. The tests are real and public; the gaps are listed on this site rather than discovered later.

Start somewhere you can afford

A staging service, or a process you own on a box you control. The guardrails exist precisely because the first attach shouldn't need courage.

Bug reports are the useful contribution

Most of what works today was fixed because an attach to real third-party software failed first. That is the highest-value thing you can send.

github.com/rajeshgangam/flux-dbg