PhysWall inverts a closed, non-linear physical law at a single measurement point — and refuses when the inverse is not unique.
The gateway between what the evidence allows and what is being claimed — a rule decides, the evidence it stands on decides how far.
Reliability is declared, never guessed from how confident a statement sounds. A certain rule standing on uncertain ground does not give a certain answer, so the weakest item here governs the whole verdict.
Each hypothesis needs a testable consequence — something that would be observed if it were true. Without one there is nothing to rule out, and the tool has no opinion worth having.
D1–D7 answer in decibels and kelvin. This answers in words, because these questions have
no instrument behind them and a percentage would be precision that nothing earned.
The rigour moves to the class of claim, and that class is stated with every answer.
Nothing here is inferred. Each hypothesis is checked against the evidence, and what contradicts
it is ruled out. Ruling out is the strong move — one contradiction is enough,
where confirming would mean ruling out everything else.
⚠ Worked out in your browser. This tool is not an inversion of a physical law, so there is no server answer to sign — it is a check on the structure of the evidence, and it says so rather than implying otherwise.
⚠ worked out in your browser
Three places, none of them ours, none of them aware this exists:
since 1962 aviation investigation facts, elimination, verdict
— and a formal "could not be determined"
1987 consistency-based diagnosis the same elimination,
written as logic rather than procedure
two analysts at NASA independently coded the same
reports and agreed — 52 contributory factors
against 55, means of 1.9 and 2.0
That last line matters more than it looks: it is the check we would have had to run before trusting any of this, and somebody ran it decades ago.
Aviation investigators have been doing this on paper since 1962. Their reports separate the facts, the reasoning, and the verdict — and the reasoning eliminates. From one report, on a fatal crash:
"No evidence of any preimpact mechanical malfunctions
or failures were observed" → mechanical, ruled out
"Neither drug is generally considered
to be impairing" → medication, ruled out
"The pilot's sleep-wake history
could not be determined" → fatigue, no verdict
→ what survived: disorientation
Three hypotheses, two eliminated by evidence, and one left open because the evidence could not settle it. That third line is the one worth noticing — a formal investigation with a legal obligation to name a cause still wrote down that it could not tell.
Their findings even carry codes: (C) for cause, (F) for contributing factor. The distinction this tool makes between one survivor and several is the distinction they already make.
Thirty of these reports have now been read this way — out of 464 distinct ones, counted from the full 1967-2021 series. The stated cause and the surviving hypothesis matched in twenty-six. In two the investigators themselves returned no verdict. And in several, the toxicology field is simply empty — because nobody was hurt, so no autopsy was done, and a whole line of enquiry closed before it opened. That is different from evidence that exists and does not decide, and this tool does not yet tell the two apart.
⚠ And counting the rest is harder than it looked. The obvious approach — search all 464 for phrases like "was ruled out" — was tested on one report and returned nothing useful: three hits on "no evidence of", all of them routine wreckage examination, and zero on "ruled out". Meanwhile the report did list four alternatives and dismiss them, under a heading that uses none of those words: "None of the following were factors in this accident." That structure is the marker, and it was found by running one report rather than 464.
And the rule above turns out to be one they mostly follow. Five of the thirty were engines that ran dry. Three name the pilot's fuel planning and put the empty tanks as the result; two name the tanks and stop. Both of those two are reports where evidence sat in the file and the verdict did not reach it.
⚠ And this is a correspondence, not a validation. Nobody has run this tool against those reports yet. What it establishes is narrower and still worth having: the shape was arrived at twice, independently, and the older one has sixty years on it.
The other tools on this site run closed physical laws backwards. This one has no law in it at all. It runs on the same four requirements anyway:
a relationship, backwards hypothesis → what you would see
becomes: what you saw → what survives
a window does this observation rule it out
a refusal no verdict when nothing survives,
and it says so when two do
a stated source every hypothesis states what it predicts,
or this will not run at all
The structure is what makes an answer trustworthy, and it does not need physics underneath to work. Physics is where it was found, not what makes it hold.
464 aviation accident reports, a public corpus. Thirty were read and their probable cause coded from the report alone, before any comparison — blind, not by consensus.
The benchmark is external too: a review of 25 published inter-rater studies in this field puts accepted agreement at 70% to 88%.
⚠ Nothing here was measured by us. The reports are public, the causes were determined by investigators who had never heard of this, and the benchmark comes from a literature review we had no part in. The only work that is ours is the coding, and it is the one thing we say plainly.
A finding from metrology applies here without change. EA-4/02's canonical worked example states an arithmetic mean of -94 nm from five readings that average -92 nm — and everything computed from that mean is correct. The arithmetic below a summary can be flawless while the summary itself does not follow from its data.
⚠ 26 of 30 is a summary too. It came from thirty individual codings, and this page prints the total without them. You cannot check it against what is shown here.
⚠ And the metrology case does not carry over in full. There, a mean of five readings either equals -94 or it does not, and one recomputation settles it. A coding decision has no such arithmetic: two readers can disagree about a report and neither be wrong, which is why this field measures agreement rather than correctness. What carries over is only the structural point — a total should be checkable against its parts.
The engine behind this tool now refuses a stated summary that does not match the values it was derived from. This page does not yet give you the values to run that check on ours — and saying so is the least we can do until the per-report codings are published.
The coding was blind: each report was read and its cause coded from the report alone, before any comparison.
A review of 25 aviation inter-rater studies puts the accepted agreement range at 70% to 88%. 26 of 30 is 86.7% — inside it, near the top.
⚠ Roughly half the published kappas in this field are post-discussion. One marine study reports raw figures of 0.45 and 0.39, then a conversation between raters, then recalculated figures of 0.72 and 0.64 — and it is the second pair that got published. 0.39 is below the usual floor; 0.64 is not.
And percent agreement is not kappa. Kappa corrects for chance; percent agreement does not. The same 86.7% gives a kappa anywhere from 0.61 to 0.82 depending only on how the categories are distributed, so this number should not be compared to a published kappa until that distribution is published — and comparing them anyway is the failure this site is about.
A checking tool, not professional advice. It tells you what a measurement does and does not support; what to do about that is your decision.