CanLab/docs

How the analysis works

All of this runs locally with no API key and nothing leaving the machine. It is also all heuristic, so this page is as much about the limits as the methods.

Read this first

These methods produce candidates. A confidence figure is a match fraction over the frames you loaded, not a statistical proof, and not a probability that the answer is right. A byte can satisfy a checksum relation by coincidence, especially in a short capture. Verify against the vehicle before you rely on anything here.

Counters and checksums

core/counter_checksum_detector.py sweeps every message.

A counter is a byte, or a nibble of one, that increments by one from frame to frame and wraps. The detector checks the whole byte and each nibble separately, because four-bit counters packed into the top or bottom half of a byte are common.

A checksum is a byte that can be computed from the others. The detector tests whether each byte is the sum, the exclusive-or, or the nibble sum of the rest.

Knowing which bytes these are matters more than it sounds: they change constantly, so they look exactly like fast-moving signals until you rule them out.

Identifying the checksum algorithm

The sweep above tells you which byte. To find out which algorithm, AUTO-RE → CHECKSUM GUESSER takes one message and one byte and tries them all: plain and inverted sums, two's complement, nibble sums, CRC-8 in its SAE J1850 and AUTOSAR forms, and the manufacturer variants used by Hyundai, Toyota, Honda and Subaru (core/checksums.py).

It fits on the first 70% of the capture and validates on the remaining 30%, then reports both numbers. That split is the useful part: an algorithm that scores well on the training portion and badly on the validation portion has memorised noise rather than found the rule.

Entropy boundaries

core/entropy_boundary.py measures per-bit entropy across a message. Bits that never change are padding. Bits that change constantly carry something. Runs of bits with similar entropy tend to belong to the same field, so the boundaries between runs suggest where one signal ends and the next begins.

It is a suggestion about structure, not meaning. It will happily draw a boundary through the middle of a 16-bit value whose high byte rarely changes.

Cross-ID correlation

core/correlation_engine.py computes Pearson r between byte pairs across different messages, aligning them to nearest timestamps and sweeping a range of lags.

This is how you find the same physical quantity reported by two ECUs, which is useful because one of them is often easier to identify than the other. A high correlation between a byte you understand and one you do not is a strong lead.

Bear in mind that with thousands of aligned samples almost anything reaches statistical significance, so judge by the size of r, not by whether a test passed.

Byte role classification

core/signal_classifier.py labels each byte COUNTER, CHECKSUM, BOOLEAN, PHYSICAL or PADDING, and ML INTEL shows the same with a confidence and the entropy behind it. It combines the detectors above with simple distribution tests: a byte with two values is a flag, a byte with a smooth distribution is more likely a physical quantity.

Anomaly detection

core/anomaly_detector.py fits a baseline from traffic you declare normal, then scores later frames against it, using a z-score per byte and an Isolation Forest over the frame vector. The use is finding the one message that behaves differently when a fault is present or a button is pressed.

core/live_watch.py runs the same baseline against a bus as it is captured. Each batch of new frames is scored in one vectorised call; a payload far from the baseline is reported once per ID per cooldown, an ID that stops arriving once until it returns, a burst when an ID arrives far faster than its fitted period, and an unknown ID once. Time is the frame clock, so a replayed log gives the same events every time. Replaying the 23-minute car log from the real-data corpus after a 60 s fit costs 1 ms per 2,000-frame batch. A short fit window flags legitimate range changes, which is what a per-byte baseline does; fit on a stretch that covers what you expect to see.

Multiplexer detection

core/mux_detector.py looks for a selector byte whose value changes which other bytes are active. Multiplexed messages otherwise look like nonsense, because the same byte offset means different things depending on the mode.

Repeated blocks

A battery pack with 96 cells does not get 96 signals in one frame. It gets a run of consecutive identifiers, each carrying the same layout for a few cells, all at the same rate and length. Per-ID analysis sees unrelated messages. core/block_detector.py sees the run: consecutive IDs (a gap of one identifier allowed) with the same DLC and a rate within tolerance, a score for how far the members agree on which bytes are constant, counters, values or noise, and a proposal for the field they share. Candidates are 8-bit bytes and 16-bit words at even offsets; the byte order is the one that reads more smoothly, and a word whose low byte never moves, or whose high byte is always zero, is dropped because the byte proposal already covers it.

One decision then becomes a candidate signal in every member. Every name ends in CANDIDATE and every description says the scale is unknown: a shared structure is evidence of a repeated layout, not of what it means. On a private EV capture of 460,024 frames the 27-message block 0x380 to 0x39A at 2 Hz is found in 0.23 s, with nine members that never change and a layout that turns out not to be uniform, which the consistency figure says plainly.

Reference calibration

core/reference_calibrate.py is the one method here that can give you a definitive answer, because it uses ground truth.

Give it an independent measurement and it searches across arbitration IDs, byte ranges and endianness for the field whose values best fit by least squares. It reports scale, offset and an R² verdict of PASS or UNCONFIRMED. core/reference_series.py reads the measurement from a CSV with a time column and any number of value columns, keeping a unit written in the header as speed (km/h), or from a GPX track, from which it derives speed by haversine distance over time, altitude, latitude and longitude. Each series is calibrated on its own.

Two clocks rarely agree: a GPS logger stamps epoch seconds while a capture may start at zero, and even on one clock a phone and an adapter drift a few seconds apart. The search first finds the lag that lines the reference up with some field in the capture, a coarse pass over binned values across the window and then a fine pass sample by sample, with ties toward the smaller lag. Overlapping wins on one ID, a 16-bit word and the byte inside it, collapse to the best reading. On the shipped sample a reference put one second ahead is found at 1.0 s and the wheel speed comes back as four 16-bit big-endian words at scale 1/32 with R² 0.9987. A periodic reference is ambiguous at lags near a multiple of its period; keep the window under half of it.

Two refinements matter in practice. Sentinel codes that mean "signal unavailable", typically all bits set, are masked so they do not wreck the fit. And a fitted scale is snapped to a neat value when doing so barely changes the decode, because a real scale is far more likely to be 0.03125 than 0.031248. Both are adapted from CSS Electronics' reverse-engineering skills.

The dialog (Tools → Calibrate signals from a reference file) runs the sweep off the GUI thread with a progress bar and a Cancel button, and the rows you pick become DBC signals as one undoable step, with the Motorola start bit written correctly for big-endian fields.

Performance

Counter and checksum detection is vectorised over NumPy arrays rather than iterating rows: on a 500,000-frame capture it went from about 183 seconds to about 2.8 seconds. The heavier analyses run in worker threads so the window stays responsive.