Behavioral Metrics · 11 min read

Conversation Baselines Before You Claim a Change

Every claim that a conversation changed is a comparison against an unstated normal. How to establish that baseline from the record itself, and when a thread is too short to have one.

The short answer

Every statement that a conversation changed contains a hidden comparison. "Contact increased sharply in April" is a claim about April relative to something, and the something is usually never stated. Where it goes unexamined, the comparison tends to be against an impression formed after the dispute began, which is the least reliable baseline available. The alternative is to establish the baseline from the record before looking at the period in question. Take a window well before the events at issue, measure the ordinary shape of the conversation over it, messages per day, direction balance, typical reply gaps and active hours, and only then compare. This costs nothing, because the same measurements are already being computed, and it converts an assertion into a comparison a reader can check. It also frequently disproves the claim it was meant to support, which is the point: finding that out during your own review is better than finding out later.

Choosing a baseline window that holds

Three properties matter. It should precede the disputed period, so it is not contaminated by the events under examination. It should be long enough that ordinary variation averages out -- weeks rather than days, since a thread measured over a single week can be dominated by one holiday or one argument. And it should be a period both participants would describe as unremarkable, which is a question worth asking the client rather than deciding from the file. The last property is the one most often skipped, and it is where baselines fail. A window chosen because it makes the later period look dramatic is not a baseline; it is the conclusion restated as a method. If the choice of window changes the result materially, that fact belongs in the summary rather than in the part nobody mentions.

What to measure on both sides

Volume per day is the obvious one and the least informative alone, because message counts follow life events that have nothing to do with a dispute. Direction balance is more stable and more revealing when it shifts: a thread that ran roughly even and became heavily one-sided has changed in a way raw volume can hide. Reply latency distributions are worth comparing as spreads rather than averages, since a shift from consistent short gaps to alternating instant replies and multi-day silences is a real change that identical means will conceal. Active hours are the fourth: messages arriving in the ordinary daytime pattern and messages arriving through the night describe different situations even at the same volume, and a shift in the hour distribution is often the clearest signal in the record.

When a thread has no usable baseline

Short records are the common case. A conversation that only exists because a dispute started has nothing to compare against, and constructing a baseline from its first two weeks is a statistical fiction. The honest output is that the record does not support a change claim, and saying so is more defensible than manufacturing a comparison. Records with a structural break are the second case. If participants changed platform, one bought a new phone, or a parenting app was introduced mid-record, then the export before and after those events measures different things. A drop in SMS volume when a co-parenting app was adopted is a migration, not a withdrawal, and a baseline drawn across that break will produce a confident and entirely wrong result. Very sparse threads are the third. Where a conversation averages under a message a day, day-level metrics become dominated by noise: a single busy afternoon can look like a step change, and a quiet fortnight can look like withdrawal when it reflects nothing at all. Below that density the appropriate response is to read the messages rather than to chart them, and to say plainly that the record was too sparse to measure.

Why this survives challenge better

A stated baseline converts a characterisation into a comparison with a method attached, and a method can be checked. "Messages per day rose from a March-to-May median of four to a June median of thirty-one, on the export hashed at intake, excluding the platform change on 2 June" is a sentence an opposing reviewer can test. "Contact escalated dramatically" is one they can only dispute. It also anticipates the strongest objection, which is that the comparison period was chosen to produce the result. Documenting the window, why it was selected, and what the figures look like under an alternative window removes that argument before it is made. Under Rule 1006 the underlying records stay available regardless; a baseline stated up front is what makes that availability useful rather than merely compliant.

What is a baseline in text message analysis?

A period of the same conversation, before the events in dispute, measured to establish what was ordinary for those participants. Volume, direction balance, reply gaps, and active hours over that window become the comparison for any later claim that something changed.

How long should a baseline period be?

Long enough that ordinary variation averages out, which usually means several weeks rather than days. A single week can be dominated by one holiday, one trip, or one argument, and a baseline that unstable will support almost any conclusion you point it at.

What if the conversation only started because of the dispute?

Then there is no baseline and the record does not support a change claim. Building one from the first weeks of a thread that exists because of the dispute measures the dispute. Reporting that the comparison cannot be made is more defensible than manufacturing one.

Does a change in message volume prove anything?

On its own, no. Volume tracks life events, platform migrations, new phones, and travel as readily as anything relevant to a dispute. A change is a reason to read the period, and it needs an explanation from the messages and circumstances before it means anything.

What if participants changed messaging platforms mid-record?

That is a structural break, and a baseline drawn across it will be wrong with confidence. A fall in SMS volume when a parenting app was adopted is a migration, not a withdrawal. Either compare within one platform or bring both records into the case.

Which measurement changes most reliably?

In practice the hour distribution and the spread of reply gaps shift more informatively than raw volume. Messages arriving through the night, or a move from consistent gaps to alternating instant replies and multi-day silences, describe a change that identical daily totals will hide.

Published by

Textimony. Editorial status: Source-linked informational guide. Updated: 2026-07-30.

Sources

Federal Rule of Evidence 1006: Summaries to Prove Content; Federal Rule of Evidence 901: Authenticating or Identifying Evidence; NIST SP 800-101r1: Guidelines on Mobile Device Forensics