A physics engine, and a verification tool enforced on top of it. PhysWall inverts a closed, published, non-linear physical law at a single measurement point — and refuses when the inverse is not unique. Seven laws, one engine, and the same refusal in all of them.
A physical-logic engine that decomposes a measurement gap along four axes — the machine, the law, the object, and which definition — and runs every domain through one structure.
RF matching, PCB loss, orbital conversion, photonic scattering, stream gauges, ocean waves, basketball. Same engine, same refusal: when the measurement cannot carry a conclusion, it says so instead of answering.
EA-4/02 is the European uncertainty guide, and its worked example S4 — a 50 mm gauge block — is the case everyone calibrates against. Two published editions give two answers:
measured value, both editions 49,999 926 mm
M:2013 ±73 nm
EA-4/02 (ENAC) ±69 nm
5.8% apart
The estimate did not move. Only the uncertainty did.
The natural way to check a tool like this is to run the canonical example and see whether it matches the published number. ⚠ There are two published numbers.
So matching one would prove nothing, and matching both is impossible. What a tool should say here is that its two most authoritative sources are inconsistent, and by how much — which is the same result the energy page reaches about water: the choice of document sets how wide the answer is, before anything is measured.
We opened both. The two editions do not use the same raw observations. The comparator readings, in nanometres:
1999 -100 -90 -80 -90 -100 newer -100 -95 -80 -95 -100
⚠ And the 1999 edition states an arithmetic mean of −94 nm from five readings that average −92 nm. Its own uncertainty budget then carries that −94 forward, and its RSS reproduces exactly — we recomputed all seven contributions and got 36.4 nm and U = 73 nm, matching the document line for line.
So the arithmetic below the mean is sound and the mean itself does not follow from the data above it. The newer edition changed the two readings to −95, which do average −94 — the observations were corrected to fit the number, rather than the number recomputed from the observations.
This is the worked example that calibration software validates against. QMSys states its own validation was performed by solving the examples from EA-4/02, UKAS M3003 and EURACHEM CG 4. A worked example whose stated mean does not follow from its own data is a strange thing for an industry to calibrate itself on.
A second published budget, from a 2025 preprint, checks out on its combined uncertainty and not on its percentages: the contributions were computed from rounded squares rather than exact ones, so a term published as 2.4% is 2.52%. Small, real, and exactly the kind of thing that survives because nobody recomputes a table they are quoting.
The five tools here take the measurements most people already have. Some organisations measure something else, continuously, and want the engine fitted to that instead.
a professional team its own shot log, its own
court, its own season
an insurer or lender claim files and loss records,
for risk work
a large laboratory its own instruments, its own
calibration chain and units
In each case the physics does not change and the fitting does. A team's release consistency, an insurer's contradiction rate and a lab's uncertainty budget are three different inversions of the same closed laws, and each needs the bands, the thresholds and the refusal conditions set against your data rather than against ours.
Run one of the free tools on your own numbers first. The gauge and the wave tools cost nothing and never will — and a message that arrives after you have put real data through one of them tells us what you measure, which is the only thing that makes a first conversation worth having.
⚠ There is no price on this page, and that is deliberate. Fitting an engine to a laboratory's calibration chain is not the same job as fitting it to a season of shot data, and quoting one number for both would mean quoting work nobody has looked at yet.
The five standard tools are priced and stay priced. This is scoped first and quoted after — and if the standard tool already answers your question, we will say so and you will not need this.
Seven claims, each with the check that would break it. Two of them were added after outside data changed what we could say, and one was rewritten because that data showed the original wording claimed too much. None needs our cooperation, and none needs anything installed beyond a browser and the papers named.
Claim: the disk model reproduces seven published
anapole wavelengths to 5.2% mean error, 15.5% worst — and one
coefficient in it was fitted to those same seven, so this is a
residual and not a prediction.
Check: the seven radii, heights, refractive indices and measured
wavelengths are printed on the bounds page with their sources. Take any
one paper, read the geometry out of it, put it into the tool, and
compare.
Breaks if: a paper says something other than what the table
says, or the tool disagrees with its own table.
Claim: ka = 1.28–2.57 across those same
seven.
Check: ka = 2πR/λ. Seven multiplications. The page
states every R and every λ.
Breaks if: your arithmetic disagrees. This page previously
claimed 1.14–1.53 and that was wrong — the correction and
the seven values that refuted it are still printed, which is what a
checkable claim looks like after it fails.
Claim: the pre-Harvey and post-Harvey curves for
one USGS gauge cross 10,000 cfs 2.53 ft apart, which is about
twice the ±8.1% uncertainty band.
Check: both curves come from USGS, which publishes them. Fit
them yourself.
Breaks if: USGS's published curves give a different crossing.
Claim: two published figures for the same
laminate differ by 33%, and both are models rather than
measurements.
Check: arXiv:1608.04347 is free. Search it for "0.37" and read
the sentence around it.
Breaks if: the paper measures what we say it expects. We
claimed it measured, and it does not — that correction is on the
bounds page.
Claim: the engine refuses rather than choosing
when an inversion has more than one answer.
Check: give any tool here a measurement that two different
inputs produce. Malus at 1° and 179° is the easy one.
Breaks if: you get a single number with no note.
Claim: EA-4/02's worked example S4 gives
±73 nm in one published edition and ±69 nm in
another, for the same measured value.
Check: both editions are public. Open them and compare the two
uncertainty lines.
Breaks if: they agree, or the difference is not 5.8%.
Claim: 26 of 30 is 86.7%, the coding was blind,
and the same 86.7% corresponds to a kappa anywhere from 0.61 to 0.82
depending on category distribution.
Check: kappa = (po − pe)/(1 − pe). Put 0.867 in with
any four-category distribution you like.
Breaks if: the spread is narrower than we say, in which case the
number is more comparable than we claimed and we were too cautious.
⚠ More than a passing one. Four outside reviewers have checked this system, and every one of them found something: a range that its own cited file contradicted, a claim of measurement where the paper said expectation, a solver converging in a regime the page itself rules out, a bundle shipped without two of its entry points.
All of those are still written down here, next to what replaced them. A site that only shows what survived is a site you cannot check — you would have nothing to compare the surviving claims against.
ASME VVUQ 1–2022 splits a computational claim three ways, and the split is worth borrowing because it is the honest description of where this stands:
verification does the code fit the mathematical
description
validation does the model represent the real
world application
UQ how do variations propagate
Verification — strong. 55 runnable checks, 518 tests , one exhaustive enumeration over every case of the search behaviour rather than a sample, and one property proved symbolically. Every number on this site reproduces from the source it cites.
Validation — one domain, and less than it sounds. D7 reproduces seven published anapole wavelengths to 5.2% mean error, worst case 15.5%. The seven were measured by seven other groups who had never heard of this.
⚠ But one coefficient in that model was fitted to those same seven points, so 5.2% is a fit residual and not a blind prediction. One free parameter against seven measurements leaves six degrees of freedom, and the residual is nearly flat across the fitted range — 5.2%, 5.1% and 5.2% at three trial values a third apart — so the agreement comes mostly from the form of the model rather than from the value that was tuned. That is worth something and it is not a prediction.
⚠ It became one. Three GaN disks from arXiv:2111.08937 — a material and a wavelength the rule never saw, at three heights that move H/R across almost the whole calibrated range — land at 4.4% mean against the coefficient as fitted. Nothing was retuned.
And the same paper publishes its own design rule for the same
dependence, D = aH^b, with both parameters fitted.
So the dependence on the height-to-radius ratio is not our
finding — this expression of it is, and it uses one
coefficient where theirs uses two.
Every other domain has zero external measurements, which is why they say `beta` and not `validated`. The test that would settle D7 is an eighth anapole, published after this coefficient was fixed, that the model was never shown.
Uncertainty quantification — this is the product. Every answer carries a band. Every inversion reports its conditioning. A non-unique inverse is refused rather than resolved by preference. And the page on energy shows the same physical uncertainty giving ×1.14 under one law and ×3.29 under another, which is a UQ result before it is anything else.
⚠ "Validated" here means one thing: an outside measurement moved a number on this site. It does not mean certified, it does not mean audited, and there is no body that certifies an inversion engine. What there is instead is that every claim can be checked — the code is readable, the sources are named, and the arithmetic reproduces.
A statement without an outside check is a statement. The seven anapoles are the outside check this system has, and one domain is what one outside check buys.
A sample reads 1% modern carbon. How old is it, and how precisely can you know?
f = 0.5 ^ (t / 5730) t in years
Reading it backwards: t = 8,267 years for that fraction. And the part worth sitting with is the precision. A 1% error in the measured fraction gives a 1.44% error in age at 5,730 years, and only 0.18% at 46,000.
⚠ This read “1% at 5,730” until an outside reviewer worked the amplification out: it is T / (ln2 · t), which equals 8,266.6 / t. That is 1.000 at 8,266.6 years and 1.443 at 5,730. The unity point is the mean lifetime, not the half-life — and 8,266.6 is the same figure already on this page as the age of the sample, attached to the wrong year.
The inversion gets better with age, not worse.
Because the derivative of a logarithm shrinks. The same measurement error buys you more precision the further back you go, right up until the signal itself runs out.
That figure is exact for pure exponential decay. Real radiocarbon does not decay into a pure exponential in calendar time — atmospheric carbon-14 has varied, so the curve from carbon age to calendar age wiggles.
intcal.org/curves/intcal20.14c 9,500 rows · plain text · no key, no registration cal BP 14C age Error Delta 14C Sigma
Uncertainty on both axes. The question above has one error term. A real sample has at least four: the counting statistics, the curve's own spread, whatever the sample did after it stopped exchanging carbon, and which version of the curve was used. Those move independently, and the clean formula sees only the first.
target periapsis at Mars 226 km minimum survivable altitude 80 km where it arrived, reconstructed 57 km signal lost 49 seconds early the spacecraft's own software computed correctly, in metric the ground software computed correctly the trajectory model computed correctly arithmetic checked afterwards no error found every figure above is from the NASA mishap investigation board report, 10 November 1999
⚠ And the 57 km is worth reading carefully: it is a trajectory reconstruction, not a measurement. Nothing measured it, because the spacecraft was destroyed. What is measured is the 49 seconds — the signal stopped earlier than it should have, and the altitude is what the models say would produce that.
Everything in that list is true and none of it is a mistake. Check the maths and it holds. Check the sensors and they were fine. Check the physics and it is textbook. So where did 169 kilometres go?
Mars Climate Orbiter, 1999. The ground software that processed thruster firings produced impulse in pound-force-seconds. The interface specification required newton-seconds. Both numbers describe the same physical event correctly. They differ by 4.44822.
The navigation model understated every thruster firing by that factor, across nine months. The spacecraft arrived 169 kilometres lower than planned, 23 kilometres below what it could survive, and burned up.
Nothing was measured wrong. Nothing was calculated wrong. The question of which definition was meant was never asked.
Which is why this site separates a gap into four things that move independently, and why one of them is which definition was used. Three of the four were faultless here. The fourth cost a spacecraft.
⚠ And this is not an argument for metric. Either unit would have worked. What failed was that a number crossed a boundary between two teams carrying no statement of what it meant. Which is why every threshold on this site names where it came from, and why a declaration that does not is refused.
Spacecraft passing Earth for a gravity assist have come away with slightly more speed than the models predicted. Millimetres per second, measured by Doppler tracking, published in Physical Review Letters, and unexplained since 1990.
Galileo 1990 +3.92 mm/s NEAR 1998 +13.46 mm/s Rosetta 2005 +1.82 mm/s Rosetta 2009 no anomaly Juno 2013 no anomaly
Every systematic error source has been modelled by JPL, Goddard and the University of Texas. None accounts for it. Proposed explanations include dark matter around Earth, modifications to general relativity, and a variable speed of light.
Not an explanation. A question about whether there is anything to explain. From the 2014 reconstruction of the Juno flyby, buried in a paper about something else:
"a high-precision gravity field of at least 50x50 coefficients was needed for accurate flyby predictions. Use of a lower-precision gravity field would yield a 4.5 mm/s velocity error."
Two of the three anomalies are smaller than that:
Rosetta 1.82 mm/s inside the modelling uncertainty Galileo 3.92 mm/s inside it NEAR 13.46 mm/s three times larger — survives
No verdict on two of them. Not because the answer is unknown — because the measurement cannot separate a real effect from a modelling choice, and no theory will fix that. Only a better measurement will.
That closes two lines of inquiry rather than leaving them open. And it leaves NEAR, which is three times too large to explain this way and is the one that still needs an answer.
⚠ What this does not claim: that the anomaly is solved, that the teams involved used a coarse gravity model — they would not have — or that nobody has said this before. The 4.5 mm/s figure is theirs, not ours. What is ours is putting it next to the anomalies and asking whether the difference is resolvable. For two of them it is not.
PhysWall was developed and architected by Gadi Zion.
Built on PhysWall — the same engine reads antenna bandwidth, conductor loss, bit erasure and heat limits. It answers what the measurement implies, and refuses when the measurement cannot say.⚠ Check this instead of believing it. Every number here reproduces from a source that is named, and the claims that turned out wrong are still printed next to what replaced them. The same engine runs all of these — it asks how much a measurement allows you to conclude, and refuses the same way in every field. The same engine runs all of these — it asks how much a measurement allows you to conclude, and refuses the same way in every field. How to check each one →