Files
Kevin Svetlitski 52b9b4863f Introduce new, experimental trace-writing backend
The existing `Trace_writer` module works well enough (albeit not perfectly) most of the
time. However, it is difficult to reason about, in large part because it writes the
trace in a streaming fashion. That introduces significant additional complexity and
bookkeeping, and limits the ability of the trace-writer to make use of information
discovered later in the trace (I believe the latter is why traces produced today often
have the few frames closest to the root wrong). Because we want to extend the trace-writer
with new functionality, we're starting fresh with a different design that's easier to reason about.
The new implementation currently exists alongside the original, but the goal is to eventually
replace it entirely.

Instead of writing the trace in a streaming fashion, we construct an internal
representation of the trace in memory, and write out the trace in a separate, final pass
once all of the events have been consumed. The module responsible for doing most of the
heavy lifting is the new `Trace_segment`, which represents a continuous, lossless, and
error-free segment of the trace; we create a new trace-segment whenever we encounter an
error.

The other major addition is that the new implementation includes inlined function calls,
using LLVM for symbolization, dramatically increasing the fidelity of the trace.

**This PR is effectively an alpha of the new implementation.** The code here does indeed work,
and produces better traces than the existing backend in many cases, but there are a couple
critical issues:
1. **Error recovery**: We create a new trace-segment whenever we encounter an error,
   **but at present we naively treat each trace-segment as disjoint**. We need to add an additional
   "stitching" pass before the trace is written out, making a heuristic, best-effort attempt to
   join together adjacent trace-segments in a way that preserves control-flow continuity.
   All of this is a long way of saying that if you encounter *any* error while using the new
   implementation, your trace is likely to be horribly broken.
2. **Performance**: Including inlined frames makes the traces significantly larger, and the supporting
   code for this new functionality is written pretty naively from a performance standpoint. As a result,
   the new implementation is roughly 2x slower than the old implementation.

It should also go without saying that while this code appears to work well on the traces I've
tried it on, I would not at all be surprised if there are still bugs/edge-cases.

We will address these shortcomings over time, but in the meantime the new implementation is opt-in;
setting the environment variable `MAGIC_TRACE_USE_NEW_TRACE_WRITER=1` will enable it.

Signed-off-by: Kevin Svetlitski <ksvetlitski@janestreet.com>
2026-05-27 10:25:46 -04:00
..
2022-01-26 20:57:02 +00:00
2022-01-26 20:57:02 +00:00
2022-03-17 10:53:42 -04:00
2022-01-26 20:57:02 +00:00
2022-01-26 20:57:02 +00:00
2022-01-26 20:57:02 +00:00
2022-01-26 20:57:02 +00:00

The magic-trace direct backend and ideas for the future

The direct_backend subdirectory of magic-trace contains an alternative backend which directly uses perf_event_open and the libipt library to capture the Processor Trace data and decode it. This is faster for decoding than going via the perf command, and gives more control, but had more gotchas and needed more work than initially expected.

It works, but currently it doesn't support decoding usage of the Linux vDSO, which means whenever a vDSO syscall like getting the current time happens, it causes a trace decoding error leading to a short gap in the trace. It's also missing many other features perf has like multi-threaded recording and capturing multiple snapshots.

It remains here because it represents a substantial investment in figuring out how to directly use perf_event_open with Intel Processor Trace and libipt together, the only such example code I know of. Should anyone want to do something that requires deeper control over Processor Trace, they may need this code. That's also why it's open-sourced despite it not (yet) being set up to build outside the Jane Street tree, because it's potentially valuable reference code.

The demand for instruction-level information

The main reason we started down the path to this backend is that the basic perf backend of magic-trace doesn't properly follow the effect of OCaml exceptions on the stack. Doing so requires noticing the instructions that push, pop and raise from the OCaml exception handler stack, which aren't branches and so aren't listed in perf's branch output.

The initial purpose of the libipt backend, as well as increased decoding performance, was enabling instruction-by-instruction decoding so we could follow exceptions properly.

An alternative approach

During the process of implementing the direct backend, it became clear that the task was more difficult than we thought, and a picture of an alternative approach emerged, on that seemed like potentially less work than bringing the direct backend up to where perf is.

Here's a sketch about what the future of magic-trace could look like:

perf dlfilter

Newer versions of perf have a feature called perf-dlfilter that allows perf to load a shared library which gets callbacks for decoded events using a C API. This was specifically designed to allow consuming Intel Processor Trace data very quickly and without text parsing.

We could implement a small shared library in something that's easy to connect with a C interface, like C/Rust/Zig, that does the following:

  • Filter out jumps that are within the same symbol, currently we need to process and discard all of these in OCaml, which is inefficient.
  • Write out info about the relevant events in an efficient binary format that's easy to consume from OCaml, potentially even Fuchsia Trace Format.

This would achieve the goal of improving decoding performance with much less tricky C code and without reimplementing so many perf features.

perf --control

Currently we tell perf to take snapshots by sending it SIGUSR2. This requires some arbitrary wait timers and has some issues reliably capturing snapshots when a process shuts down. The direct backend's use of perf_event_open was going to solve this.

However newer versions of perf add a --control flag that allows using a FIFO with acks to toggle events and take snapshots, removing the need for signals and waits, and also a --snapshot=e option to guarantee a snapshot on end if there hasn't been another snapshot.

instruction-level decoding using basic blocks

In order to handle OCaml exceptions we need instruction-level decoding. This could've been done with libipt but an alternative way is to use a separate instruction decoding library/tool to decode the instructions between branch events provided by perf.

This would look something like having a cache of computed info for each previously encountered basic block, for example the exception handler pushes and pops present in it. When a new basic block (start and end address of execution with no intervening branches) is encountered, that range of the binary is decoded with a disassembler library and the necessary info computed from it and put in the cache.

There's a few potential ways to do this:

  • Do the basic block decoding inside the perf-dlfilter stub and then pass out the decoded info.
  • Use an instruction decoding library from OCaml.
  • Use an offline disassembler tool like objdump on the entire binary and fetch basic blocks from that and parse them from text. This is probably the simplest to do from OCaml but would run slowest.