← Back to articles

MEASUREMENT · V&V · AUTOMOTIVE QUALITY

18 min read

When can you trust a measurement? Notes from validating an algorithm on real cars.

An algorithm that returns a number is not finished. It is finished when you can state the conditions under which the number is wrong — and test for them.

My thesis work sits on a specific problem: computing GAP and FLUSH between the body panels of a car — the width of the seam between a door and a wing, and the height difference across it. They decide whether a vehicle looks right and whether it seals against water and noise, and they are measured on the line with a laser profilometer.

The algorithmic part is geometry. The engineering part is a different question entirely: how do you know the number is right?

The framing that changed how I work: a measurement algorithm is a claim about a physical object. Testing it is not checking that the code runs. It is establishing the conditions under which the claim holds, and demonstrating that you detect it when they don’t.

1. There is no ground truth, and that is the whole problem

In ordinary software testing you write an expected result. Here nobody knows the true gap to a hundredth of a millimetre — that is why you are measuring it. This is the oracle problem in its purest form, and it forces a different approach.

What you can do instead:

  • Compare against an independent method. Two algorithms disagreeing tells you something even when neither is known correct.
  • Assert properties rather than values. Reverse the profile and the gap must not change. Translate it and the gap must not change. Scale it and the gap must scale exactly.
  • Use synthetic profiles with a constructed answer. If you build the geometry, you know the truth — and you can sweep it across the whole parameter space.
  • Compare against a slow, obviously-correct reference implementation. Too slow for production, ideal as an oracle.

2. The precondition is the specification

Every geometric method assumes something: that two edges are identifiable, that the surfaces either side are locally planar, that the profile actually contains a gap. When those assumptions hold, the method works. When they don’t, it returns a number anyway.

# The dangerous version
def gap(profile):
    left, right = find_edges(profile)
    return distance(left, right)      # always returns something

# The version that can be trusted
def gap(profile):
    left, right = find_edges(profile)
    if left is None or right is None:
        return Result(None, reason='edges not identifiable')
    if planarity_error(profile, left, right) > TOL:
        return Result(None, reason='surface not locally planar')
    return Result(distance(left, right), reason=None)

The second version is longer and it is the only one that belongs on a production line. A refusal is information. A confident wrong number is a defect that reaches the customer as a mis-fitted door.

3. Test cases come in three families

FamilyWhat it establishes
Nominal — clean profiles, geometry well within the envelopeThe method computes correctly when everything is as assumed
Boundary — at the edge of the validity domainWhere it stops working, and whether it says so
Known problem cases — profiles that previously produced bad resultsRegression: the failure does not return

The third family is the one that accumulates value. Every profile that produced a wrong answer becomes a permanent case. After a few months the suite is not a generic benchmark — it is a record of every way this specific problem has gone wrong.

4. The tool needs testing too

I built a desktop environment to run these comparisons profile by profile and show the geometry behind each result. It is easy to treat that as scaffolding rather than software, and that is a mistake: if the visualisation misplaces an edge, or the comparison pairs the wrong profiles, the conclusions are wrong even when every algorithm is correct.

So the tool carries unit tests, regression baselines on stored outputs, and edge cases of its own — empty input, a single point, a profile with no gap at all. Analysis code decides verdicts, which makes it production code.

The failure mode nobody watches for: an analysis script that averages away a discrepancy instead of reporting it. Every difference above the acceptance criterion should be accounted for individually. The moment you start explaining differences as noise without evidence, the validation has stopped being validation.

5. What transfers

None of this is specific to body panels. The same structure applies to any measurement a decision depends on:

  • State the preconditions and the validity domain, in writing, before testing.
  • Agree the acceptance criterion before you look at the results — afterwards it becomes whatever you achieved.
  • Prefer a refusal to a wrong answer, and make refusals visible.
  • Keep every real failure as a permanent test case.
  • Treat the analysis tooling as seriously as the algorithm.

It is the same discipline as software testing, pointed at physics instead of at a web application. The question does not change: not does it produce a number, but under what conditions is that number wrong, and would I know?

When Can You Trust a Measurement? — Ibrahim Kenia