Does this actually work?
For someone with no reason to trust this and no time. Every figure below names the document in the repository that owns it, so you can check the owner rather than this page.
The number to start with
Where this engine computes a date or a float, it matches what Primavera P6 stored on 67.85% of 584,687 field comparisons, across 82 real projects and 86,031 activities, P6 5.0 to 24.12. Over the 66 projects whose stored dates reconcile against their own declared calendars it is 78.06%, and the sixteen projects the validity gate refuses hold two fifths of all the comparisons — so that second figure goes beside the first and never instead of it. The number went down, not up, when the corpus grew.
Owner: engine/oracle/corpus/SWEEP-2.md.
That is close to one field comparison in three where this engine disagrees with the software that produced the file. It is measured, not estimated, and it is not a number anyone should be comfortable with.
And almost none of that gap is Primavera's
Primavera was measured against itself over the same 584,687 comparisons, with
this engine's scheduler never called. 11,330 of them sit on a row whose stored
values cannot all be true at once, so achievable agreement is capped at
98.06%. The apportionment of the remaining shortfall between the two
instruments is published by engine/oracle/corpus/P6-SELF.md, which owns it —
and it is overwhelmingly ours, not theirs.
That figure is here to close the excuse, not to soften the number. It is the opposite of a mitigation.
It fell, and the movements are not one movement
The headline has been 79.32%, then 60.51%, then 61.77%, 61.89%, 67.50%, 67.54%, 67.71%, 67.80%, 67.86% and now 67.85%. None of those earlier figures is this engine's agreement and they must not be quoted as such.
More importantly: a page presenting them as one number improving nine times is telling a false story out of correct digits. No two of those movements are the same kind of movement.
- The drop from 79.32% happened because the corpus grew fifteenfold — the older population was mostly clean baselines, the new one carries progressed, multi-calendar work.
- One rise came from two arithmetic fixes in the scheduler.
- The largest single rise came from two defects in the importer, not the scheduler.
- One rise was a corrected definition rather than more agreement: Primavera's driving-path flag marks membership of a set of driving relationships and this engine had been publishing a single chain, which is a subset of that set. The same evidence, scored against a corrected question.
- The most recent movement, on 6 September 2026, went down: 27 comparisons in two projects, both falling, with the other eighty projects holding cell for cell. It was kept and reported falling, because the reading behind it is right.
Two figures keep the whole thing in proportion. Over the fifteenfold expansion the per-project median went down, 85.7% to 85.6%, while the aggregate rose. And a single arithmetic defect was worth 1.26 points on its own — which says how much was still wrong, not how far this has come.
There is a standing rule in this project and it is the discipline the entire validation claim rests on: agreement is a measurement, and the moment it becomes a target it stops measuring. There is precedent in both directions. A fix that lowered agreement was kept because it was right. A plausible fix that would have raised it was falsified and refused.
The other measurements
Importer fidelity. 1,923,733 of 1,947,919 field comparisons — 98.76% —
against MPXJ over 80 projects in 65 files, every disagreement carrying a named
cause. It measures agreement with another importer's convention, not with P6's
stored dates. Owner: engine/oracle/FIDELITY.md.
Rule dependence. 265 of the 351 rules never touch the CPM arithmetic, 56 rest
on it, and 30 are engine-gated. The three sum to 351 and a test asserts that they
do. For UFGS specifically, 35 of 48 rules are independent of the engine entirely.
Owner: docs/RULE-DEPENDENCE.md. This is the most useful number on the page
if you are evaluating the conformance layer, because that layer is not capped
by 67.85%; it is capped by importer fidelity.
Mutation score. 1,553 of 1,742 mutants killed, 89.2%, over twenty modules
at a stated revision and seed. Read the spread, not the aggregate: one module at
64.8%, four at 100%, and eight modules never measured at all. Four earlier scores
in the high nineties were withdrawn when the harness was caught reporting kills it
had not made. Owner: engine/quality/MUTATION-REPORT.md, which states that an
aggregate "should not be quoted on its own".
Verdict yield. Just under two fifths of rule-file pairs return a verdict on a
bare run; a little under half when the submission names its baseline; and about
three fifths on the US federal path once the contract's own terms are supplied.
Every one of the twelve-thousand-odd abstentions names a specific missing input.
The exact figures are owned by engine/quality/VERDICT-YIELD.md §14 and move
often enough that they are deliberately not restated here.
That number has fallen three times, and each fall was an improvement. One rule had published a pass sixty-six times on a condition it could not have failed for any input whatsoever; it was withdrawn. Five corpus files holding several projects each are now refused outright rather than reviewed as whichever project happened to be written first — one of them had been producing a 181-clause federal report about a one-activity stub out of 886.
Answering is not discriminating. Every contract input a reviewer can supply adds many rules that answer and far fewer that tell one schedule from another. A higher answered count is not automatically a better product, and this project measures both.
Performance. 19,202 activities and 28.6MB in 13.5 seconds on a 2017 four-core
mobile i7. Owner: engine/quality/PERFORMANCE.md.
The claim this project refuses to make
There is a cross-validation fixture in the repository — 27 activities across 13 networks, captured from a real Primavera P6 23.12 installation — on which this engine agrees with the stored answers on 156 of 160 field comparisons.
That is the weakest evidence here and it reads as the strongest, so it is
stated with its caveat or not at all. The file's ERMHDR record names the
exporting tool as danaf, not Primavera, so the stored dates it is checked
against may not be P6's. Its honest use is as a regression check on 27
activities, not as evidence about Primavera. None of the thirteen networks is
wider than a two- or three-activity chain, and the correct reading of the
percentage is "no divergence detected on the semantics these thirteen cases
exercise" — never "agrees with P6".
Similarly, this site does not say "validated against Primavera". The repository's own wording is "validated against Primavera — partially, and measured", and the qualifying two words are load-bearing.
What has never been run
Ask this before the reproduction question rather than after it.
Continuous integration has been dead since 3 September 2026, on a billing state, with hundreds of commits landed since. Of the two hundred most recent workflow runs, none concluded successfully.
- The reproduction matrix has never executed once in the life of the repository. Cross-version determinism is unknown, not established. Cross-platform determinism is assumed.
- What is measured is an identical answer digest across five fresh processes under five explicit hash seeds — on one machine, one operating system and one interpreter. Five seeds on one Windows box is not two platforms.
- The three-way differential against MPXJ has never run automatically either.
None of that means the engine is non-deterministic, or broken on Python 3.12, or unusable on Linux. It means nobody knows, and the difference between unmeasured and fine is the whole subject of this section. A check that has never run is not weaker evidence than a check that has. It is not evidence.
What is not measured at all
- Nobody has run this on a real project. Zero customers, zero users, zero pilots, zero agencies. No professional engineer has reviewed or adopted any output.
- The customer's actual job is a monthly sequence and the corpus cannot test one. Six genuine schedule series were found in the corpus and every one of them is exactly two data dates long. There is no series of three or more this engine can load. It answers "can this be used month after month?" with "not yet."
- Every claim about what a reviewer wants comes from documents — specifications, standards, published research and one InTrans/FHWA study of eleven state DOTs — and not from reviewers.
Walk away if
- You need a validated recomputation engine. 67.85% is not a qualification for work an opposing expert will cross-examine, part of the remaining gap has a measured ceiling rather than a fix date, and no customer has ever run this.
- You need an installer, a user interface, or a certification. There is a command line and a Python library. There is no published package. There is no SOC 2, ISO 27001 or FedRAMP.
Look further if
- What you need is the conformance layer. 265 of the 351 rules read the file rather than the arithmetic. That layer is not capped by 67.85%.
- The discipline is what you are buying. The evidence for it is behavioural rather than asserted: a leak whose removal dropped agreement was removed and the fall published; four mutation scores were withdrawn when the harness was caught overreporting; a plausible fix that would have raised the headline was falsified and refused; and rules that could not fail for any input were found and withdrawn rather than left publishing passes. It is real and it is not fast. It is the only real differentiator here.
Check it yourself
From engine/, after pip install -e ".[dev]":
python check.py
python oracle/harness.py oracle/cases/p6-23.12-xval.oracle.xer
pytest oracle/ && ls oracle/cases/*.json
pytest tests/test_rule_dependence.py
python quality/ledger.py --markdown
And the corpus figure, which fetches by digest rather than shipping files:
cd oracle/corpus && python fetch_corpus.py --verify && python gapshape.py
Source: web/pages/is-this-real.md. Source commit date: 2026-09-09.