Limitations

Everything this tool cannot do, or has not proved. Consolidated here so that no user is surprised, and so that this page can be pointed at.

A company selling evidentiary rigour that overstates its own validation has destroyed the only thing it sells, and would be found out by exactly the audience it was trying to impress. So the position is stated first and in full.


1. The engine is not validated at scale

Updated 1 September 2026. This section previously said no held-out corpus paired real schedules with P6's own computed answers, and that the 27-activity fixture below was the only real P6 evidence held. Both statements were true when written and are now false: engine/oracle/corpus/ holds eleven harvested P6 exports, and agreement across them has been measured. The measurement is worse than the fixture's, which is why the correction is here rather than in a changelog.

The ceiling on this measurement has itself been measured, and it does not excuse us. P6 was run against P6 over the same 584,687 comparisons with our scheduler never called: 11,330 of them sit on a row whose stored values cannot all be true at once, so achievable agreement on this corpus is capped at 98.06% (engine/oracle/corpus/P6-SELF.md). Of the 32.14 points short of 100%, at most 1.73 are P6's and at least 30.41 are ours — the apportionment was tightened against us on 3 September, when joining the two instruments per cell showed that 1,218 of the 11,330 unreachable cells are ones we already agree with and so are not recoverable ground. That is the sentence to read the rest of this section through — a reader who sees the level below without it cannot tell an engine carrying thirty points of defects from one measured against a noisy target, and the answer is that the target is not noisy. The self-inconsistency does not even fall where our gap falls, and every self-inconsistent project already fails the validity gate.

Agreement with P6 is 67.85% — 396,735 of 584,687 field comparisons across 82 projects and 86,031 activities, P6 5.0 to 24.12 (engine/oracle/corpus/SWEEP-2.md, which measures and owns this figure; this page cites it rather than asserting the level as its own). The per-project table in that document is the honest presentation. Over the 66 projects that pass the validity gate the figure is 78.06%, and that belongs beside the headline rather than instead of it: the 16 projects the gate refuses hold 40.9% of all comparisons, and a harness that improves its number by choosing what to measure has stopped measuring. The gate was itself measured on 1 September: its duration-against-span check had reached one corpus project's calendar on 0 of its 770 unstarted rows and recorded a PASS anyway, because the harness read the week as a union of shift shapes and summed overlapping ones into a fifteen-hour day on a ten-hour calendar. Reading the per-weekday shifts instead takes it to 770 of 770; 494 of the 507 misfits were the instrument, 14 are real, and the file moves from an unearned PASS to a FAIL. It leaves the gate-passing population, taking it from 65 to 64, and because it agreed on only 40.5% of its comparisons that raises the gate-passing figure by 0.63 points — arithmetic, not improvement. The falsification was run before the percentage precisely because the movement flatters us: 494 of the 507 union misfits were overshoots by exact multiples of 2.5 hours, the 14 survivors are undershoots of one to eight hours in one contiguous sub-network, and a residual matching the same signature would have meant the fix was tuning.

This paragraph said 493 and "all 507" until 2026-09-02, and the arithmetic that looks like it justifies 493 is a trap. 507 − 14 = 493 assumes the 14 residuals are a subset of the 507 union misfits. They are not. engine/oracle/harness.py:1072-1080 states the split: 13 of the 14 already missed under the union, by the amounts that were not multiples of 2.5, and the fourteenth was scored as reconciling exactly by the union model and does not reconcile — a right answer for no reason at the scale of a single row, and the reason a corrected instrument is allowed to find more faults than the one it replaced. So the 507 partition as 494 instrument plus 13 real, and the residual is 13 + 1 = 14. engine/oracle/corpus/SWEEP-2.md:186-192 had it right and this page had it wrong, which is the reverse of the usual direction here. The gate now reports PASS 64, UNPROVEN 2, FAIL 16 — it read PASS 62 / UNPROVEN 2 / FAIL 18 until the base-calendar fix of 2 September took two files from FAIL to PASS — where UNPROVEN is a strict subset of the 66 passing rather than a third population (engine/oracle/corpus/SWEEP-2.md, engine/oracle/harness.py). The scheduling logic is internally consistent, deterministic and property-checked, and that is not the same thing as agreeing with the source software. Figures should be reconciled against P6 before they are relied on.

By the end of 1 September 2026 it had moved three times, in opposite directions and for opposite reasons, and they must not be run together; four further movements followed on 2 September and are set out below. It was 79.32% on a population of 22 projects and fell to 60.51% when the corpus grew fifteenfold: that fall is the correct result and not a regression, because the old population contained not one mid-project update and so measured the engine only on the easiest schedules P6 produces. It then rose to 61.77% on that same fixed corpus when a real arithmetic defect was fixed — a start-to-start lag charged twice where the predecessor had already started, which P6 has already spent against the actual start. Comparison, project and activity counts are identical either side of it; 7,392 comparisons came right.

It rose again, from the superseded 61.77% to 61.89%, on a second fix at 9188ad8. xer.py converted every duration through a calendar's declared day_hr_cnt, and eight CALENDAR rows in five corpus files contradict it — an eight-hour day on a calendar running 07:00 to 17:00. Five of the eight carry activities and those five are four distinct calendars, one appearing in two projects; the other three carry none. Every duration on the affected calendars was stretched by a quarter.

Two separate engine defects moved this figure on 1 September 2026, in two movements: 60.51% to 61.77%, then 61.77% to 61.89%. Three figures are two movements, and this page said "three defects" for a day because a third commit of the same date — f5b4fb7, which floors unstarted work at a project planned start later than the data date — was read into the chain. Its own message states what it moved: the eleven-file corpus, 55.11% to 79.32%, before the fifteenfold expansion, with the cross-validation fixture unmoved at 156/160. It never touched the 82-project figure, and SWEEP-2.md names the two that did. The reading does not change with either of them. Defects of that size sitting in a lag rule and a calendar's declared day length say how much remains, not how far this has come, and on that day 38.11% of 584,687 comparisons still disagreed. All three of 60.51%, 61.77% and 61.89% are superseded as headlines and must not be quoted as this engine's agreement; the 79.32% in that commit message is superseded too, and was never a figure for this population at all.

It rose again the next day, from 61.89% to 67.50%, and that movement was larger than every scheduler fix before it put together. Neither defect was in the scheduler. _read_calendars never followed CALENDAR.base_clndr_id, so a project calendar did not inherit its base calendar's exception days — worth 66.73% alone, moving exactly two projects, one of which went from 10,882 to 38,904 of 39,256 and from 1,304 misfits on its duration-against-span check to 0. And P6's midnight act_end_date is a day start, not a day end — worth 62.66% alone, moving exactly the seven projects WHOLEDAY.md had named in advance. Both are fields P6 stores; neither is a P6 answer read back in. Every per-project prediction registered before the fixes held to the digit.

Read it with the number that fell. Importer fidelity arm 1 — agreement with MPXJ, an independent importer — went from 99.60% to 99.43% over the same commit, and it reconciles exactly: 1,408 rows of a new ACTUAL_FINISH class where MPXJ keeps P6's raw midnight stamp and we now read it as a day start, plus 45 inherited exception days MPXJ does not follow. Arm 1 measures agreement with another importer's convention, not with P6's stored dates, so it is not a measure of correctness and a fall in it beside a rise against P6's own arithmetic is the expected shape (engine/oracle/FIDELITY.md).

It moved twice more the same day, and only one of those two is agreement. Every level named in this paragraph is superseded and is here as history: these are the movements, not the current pair, which is at the top of this page. 422cc25 took it to 67.54% by reading the second of P6's two constraint slots in the forward pass — +243 comparisons on one project, arithmetic. The move that took the headline from 67.54% to 67.71% is not that, and the matching move of 77.60% to 77.85% over the projects that pass the gate is not either — and both of those levels are superseded, the gate-passing one having since gone 77.85% to 78.06% in three further steps set out below. P6's driving_path_flag marks membership of the driving set — the backward closure over driving relationships, which branches at every tie — and is_longest_path had been marking membership of a chain, one predecessor per hop. A chain is a subset of the set it lives in, so the disagreement could only ever run one way, and it did: 2,564 rows to 94. Publishing the set drops the DRIVING_PATH disagreement class from 2,658 to 1,628 while every other class is unchanged count for count. The same evidence, scored against a corrected reading of the question — not more agreement on the same one, and this page says so because a sentence reading "agreement rose" would describe a measurement that did not happen.

Both figures moved once more later the same day, 67.71% to 67.80% and 77.85% to 77.97%, for the reason the next paragraph gives — and once more again on 2026-09-04, 67.80% to 67.86% and 77.97% to 78.07%, on a start-to-finish lag advanced in start space and read back as a finish boundary, which costs a working day whenever the day before the advanced instant is not worked. That fifth movement is the first that is arithmetic in the scheduling kernel; P6 and MPXJ both give the answer we now give on the activity that names it. And once more on 2026-09-06, and it is the first that goes down: 67.86% to 67.85% all-projects and 78.07% to 78.06% over the gate-passing projects, when a finish-side constraint date was read as the boundary it names rather than as the start instant of its day. 27 comparisons in two projects, both falling, with the other 80 holding cell for cell. It was kept and reported falling because BOUNDARY-SEMANTICS.md settles the reading from P6's own stored bytes.

A fourth movement followed, later the same day, and it is a fourth kind of thing again. c0c69ef had been reading a stored constraint stamp with a start rule where the constraint needed a finish rule — the same instant, resolved to the wrong end of the working day. Reading it on the correct day took the headline from 67.71% to 67.80%, which is 526 comparisons across all 82 projects. Over the 66 projects that pass the validity gate the same fix moved 77.85% to 77.97%, on 409 comparisons — a level since superseded in turn, so that the chain over those projects reads 77.85% to 78.06% across the three movements since. Every one of the 526 lands in a class carrying a mid-day stamp or none at all, and all three whole-day classes are unchanged to the row — which is the check on the mechanism rather than on the number. One project fell by exactly one comparison and was kept, two of its rows having moved onto P6's own stored answer; a correct fix that lowers a number is reported falling, not reversed.

Four movements on 2 September 2026, and no two of them the same kind of thing: two importer defects, then arithmetic, then a change of definition, then a stamp read on the wrong day. Six movements in two days, counting the two of 1 September above. A page presenting these as one number improving six times is telling a false story out of correct digits, and no check in this repository can catch that — which is why the chain is written out here in sentences rather than left as a row of figures.

A single arithmetic defect was worth 1.26 points. That is not a milestone. It says how much is still wrong, not how far this has come, and close to a third of 584,687 comparisons still disagree. Two figures keep it in proportion: the SWEEP-2.md sweep-population rows, 80.0% and 79.32% — superseded as headlines, still correct for their own narrower populations — did not move at all, because none of those files carries the feature — an independent check that the fix touched nine projects and not the corpus; and the per-project median went 85.7% to 85.6%, falling while the aggregate rose, because a median does not feel a fix that lands in nine projects out of 82.

A named architectural cause is on record and is not a bug: the time axis is whole-day, and an in-progress activity with a fractional remaining duration finishes at a mid-day instant in P6 that this engine advances to the next whole working day. engine/oracle/corpus/APPORTIONMENT.md measures that at 31.3% of the gate-passing gap (23,830 of 76,081), from its class table re-run on 2026-09-02 over the 66 that now pass; that document's two SUBDAY sub-analyses are marked pre-fix and are not quoted here.

Two things about that figure, and the second is a coincidence a reader will otherwise misread.

First, ~~30.9% (23,877 of 77,376)~~ is superseded. It was correct against the gap as it stood before the constraint clock landed, and the owning document has since been re-run against a frozen tree -- reproducing all eighteen of its previously published class counts cell for cell before publishing any new one.

Second, this paragraph read 31.3% once before, and that was a different measurement. It moved to 30.9% when the class table's own instrument was corrected earlier the same day: gapshape.day_bounds collapsed a calendar's whole definition to a single boundary pair, so a stamp that is a day end on its own day block read as mid-day. 4,202 rows changed column and none changed verdict, which is why that figure moved while no agreement figure did. The correction has not been reverted -- the number is back at 31.3% because the denominator fell from 77,376 to 76,081, not because the classifier changed back. Two different measurements agreeing to a decimal place is the kind of coincidence that makes a page look stable while it is being rewritten underneath, which is why it is spelled out rather than left as a digit.

And the count fell while the share rose. Sub-day lost 47 rows, 23,877 to 23,830. Its share went up because the gap fell faster. Any sentence reading the rising share as this class growing describes a measurement nobody made -- and would pass every digit-comparing check in this repository.

The shares moved with the importer fixes and the ranking moved with them. WHOLEDAY_BEYOND_ONE_DAY went from 44,915 rows to 23,551, the gap fell from 222,803 to 190,046, and the population passing the gate went from 64 to 66 — so every share on this page is a share of a smaller gap over a different population. Aggregated by family, the mid-day classes went from 44.9% to 51.9% of the gap and the whole-day classes from 27.6% to 19.3%: the fix that raised the headline made the mid-day family a larger share of what is left.

That class table is an 08cc4e4 measurement and has not been re-run since the definition change or the constraint clock. Its DRIVING_PATH row of 2,658 is now 1,628, and the constraint clock then moved four more rows — SUBDAY by 141, MIDDAY_BEYOND_ONE_DAY by 101, TOTAL_FLOAT by 247 and FREE_FLOAT by 37 — so five of the nine counts have moved since 08cc4e4, and the 190,046 total is now 188,247. Read that span carefully: across the constraint clock alone it is four classes that moved and five that held, and it is only cumulatively from 08cc4e4 that five of nine have. SWEEP-2.md puts sub-day's share of the gate-passing gap at 31.3% on the live figures — 23,830 rows of 76,081, where this page previously read 23,877 of 76,490. The share rose while sub-day's own count fell by 47 rows, because the gap it is a share of fell faster. That is the arithmetic of the denominator, not sub-day growing, and a sentence reading it as growth would pass any check that compares digits.

A class here can grow while the gap shrinks, and SUBDAY has now done it twice — 39,614 rows to 44,645 on 1 September, and 44,645 to 45,230 on 2 September while the gap fell by 32,757. A row corrected from many working days out lands one day out rather than vanishing, so it moves into the sub-day class. That reads as a regression to anyone not told why, and it is not one.

No class is largest by a margin worth publishing, and this page will not say one is. Across all 82 projects the top three are SUBDAY 23.7%, TOTAL_FLOAT 23.4% and MIDDAY_BEYOND_ONE_DAY 22.0% — 44,526, 44,016 and 41,457 rows in a gap of 188,247. The first two are 510 comparisons apart, which is less than the 526 a single fix moved that same day; the third sits 3,069 behind the first. The figure this page carried before, "a spread of 310", was the spread of the top two on the superseded counts and was never the spread of the three named beside it. APPORTIONMENT.md puts it as a tie and says in terms that any sentence of the form "X is the largest class" is a claim about a coin flip. The scale that makes that true: a single importer fix moved WHOLEDAY_BEYOND_ONE_DAY by 21,364 rows, and a correction to the classifier -- not to the engine -- moved 4,202 and swapped the second and third places recorded here. Over the 66 projects that pass, SUBDAY holds 31.3% of the gap (~~30.9%~~ superseded by the 2026-09-02 re-run of the owning document).

The second place has now swapped again, and it is worth saying how narrowly. TOTAL_FLOAT is at 24.5% and MIDDAY_BEYOND_ONE_DAY at 24.4% — over the gate-passing population that is 46 rows apart in a gap of 76,081, a tenth of a point. This page previously paired SUBDAY with MIDDAY_BEYOND_ONE_DAY as the top two; both of those digits were individually correct and the pairing had already gone stale. Two classes that have reordered twice in one day, on changes that moved no verdict, are not a first and a second. Do not quote a second-largest class from this page, and treat any sentence naming one as a claim about a coin flip — which is what the owning document says in terms.

The earlier version of this paragraph expressed that scale as a multiple of the spread. It is stated in rows now: the spread it divided by has itself moved twice, and recomputing the multiple would publish an arithmetic nobody measured -- which is the failure this page exists to describe, committed on the page describing it. Earlier versions of this paragraph named a largest class three times and were wrong twice. MIDDAY_BEYOND_ONE_DAY's mechanism is established as of 1 September 2026, which it was not when this paragraph was written.

The mid-day handoff. When P6 finishes a predecessor part-way through a day it starts the finish-to-start successor at that same instant on that same calendar date — both activities run on 16 April. A whole-day engine cannot: if the predecessor occupies the 17th, the successor starts the 18th. The successor therefore loses a second working day at the link, on top of the first lost within the activity by rounding its finish up.

That result is controlled, which is what makes it causal rather than correlational. Taking only activities all of whose predecessors' finishes already agree with P6, so nothing is inherited: 5,441 activities in 22 files hand off at a mid-day instant, and the successor's start is one working day late on 5,429 of them — 99.8%. The matched control is the identical topology at a shift end: 1,974 activities in 27 files hand off at a day boundary, and 1,944 of them — 98.5% — agree exactly. The timestamp is doing the work, not the shape of the network (engine/oracle/corpus/MIDDAY-MECHANISM.md).

So the "sub-day costs at most one working day per activity" bound is false. The replacement is bounded rather than open-ended, and it is measured rather than controlled — a carry model over 50,707 disagreeing unstarted rows, not the control population's own result:

A whole-day time axis costs at most one working day per activity in 97.8% of cases and at most two in 99.4%, and the second day arises at the link rather than within the activity.

Read the two weights differently. The 5,429-against-1,944 result is observed with a control. The 97.8 / 99.4 figures are modelled over the corpus. And the residual — 306 activities, 0.6%, where the model predicts three working days and the median is fifteen — is explicitly not claimed: the carry model reaches those rows without explaining them.

This is not a defect awaiting a patch. A pass that set out to write the fix measured the input it needs and refused it (MIDDAY-MECHANISM.md, "The bit cannot be derived", reproducible by python derive.py):

This share of the gap is a property of the whole-day time axis rather than a defect awaiting a patch: the quantity that drives it is sub-day, it is inherited across activities rather than generated at them, and reconstructing it from whole-day inputs recovers at most 70% of the affected rows while misfiring on 14% of the rows currently correct.

Each clause is measured. Inherited, not generated: on the document's own named case the offset arrives across eleven hops whose durations are every one a whole multiple of the eight-hour day — 144, 96, 48, 24, 24, 56, 56, 56, 120, 32, 24 — from an in-progress activity with 11.2 remaining hours; across the file, only 1,093 of 8,962 mid-day handoffs — 12.2% — have a fractionally-durated immediate predecessor, so seven of every eight inherit it from further upstream. At most 70%: a derivation built deliberately stronger than any engine could be — carrying the offset exactly in minutes, and choosing the driving predecessor using P6's own stored dates, an advantage this engine does not have — fires on 3,807 of the 5,441 mid-day rows, 70.0%. Misfiring on 14%: the same derivation fires on 280 of the 1,974 day-boundary control rows, 14.2%, breaking about 280 of the 1,944 starts we currently get right.

That defeats the fix as specified. It is not a proof that no derivation exists, and the document does not claim one.

None of this is good news, and it should not be read as any. No agreement figure moved for this reason: the mid-day work changed no behaviour, and the corpus documents say in terms that nothing in it disturbs the apportionment. The headline did move afterwards, from the now-superseded 60.51% to 61.89%, and it moved because of an unrelated start-to-start lag fix — not because anything here was applied. Two things moved in opposite directions in the same pass: a hedge became a bounded statement with a control behind it, and the prospect of fixing it receded from "specified, pending" to "defeated as specified". Nor does the mechanism explain the class — even in the file where it was found, the carry accounts for the early-start error on only 1,758 of 11,368 unstarted activities, 15.5%, and the largest single unexplained thing in the corpus is now 230 forward-pass origins in that file whose predecessors all agree. Nothing here should be read as a claim that the largest part of the gap is understood.

Every report the tool produces prints this limitation. Do not remove it from a deliverable.

What the validation corpus actually is

engine/oracle/ holds 27 activities across 13 networks, scheduled by a real Primavera P6 23.12 installation. It was captured from P6's own database after a human pressed F9, by the author of cpp-cpm-engine (MIT, Dana Fitkowski). That repository discloses the capture as fitted to its own engine; this engine has never seen it, so for us it is genuinely held out. It is the only first-party P6 capture held; the eleven harvested exports above are real P6 files but their stored dates were not watched being computed.

Current agreement on the fixture: 156 of 160 field comparisons, 97.5%.

The correct reading of 97.5% is "no divergence detected on the semantics these 13 cases exercise" — never "agrees with P6". None of the 13 networks is wider than a two- or three-activity chain. A clean sweep there is worth far less than a clean sweep on one real 800-activity update, and the corpus that would give the latter does not exist. engine/oracle/CORPUS-ACQUISITION.md is the plan to build it.

All four remaining discrepancies are one representational difference that will never agree: P6 stores a started activity's actual date in its computed start columns, and the harness classifies them PROGRESS_ACTUAL_AS_EARLY (engine/oracle/README.md, "The four divergences"). The free-float case on a start-to-finish tie that this section previously listed as a fifth is now closed: case 04 agrees 12/12, and engine/oracle/README.md Finding 3 records that it was closed by the rework of the forward/backward bound arithmetic in cpm.py — and that the reasoning behind that change has not been audited there. Anyone quoting 156 of 160 should know one of those points rests on an unaudited change. While the divergence stood it was deliberately not patched, because the semantics were not understood and tuning an engine to raise an agreement rate is how the rate stops meaning anything. Agreement is a measurement; the moment it becomes a target it stops measuring.

The harness itself carries a caveat worth reading: the XER's ERMHDR names the exporting tool as danaf, not Primavera, so the stored dates may not be P6's. The provenance is credible but it is not first-party.

What the corpus has already bought

One real defect, found because P6 disagreed. The backward pass and the free-float loop both took bounds from completed successors, whose late dates had been set to their actual dates — manufacturing large negative float on the predecessor of every out-of-sequence completion. Negative float is the criticality finding, so the engine was fabricating criticality from evidence that said the opposite.

Every property test held while that bug was live. The arithmetic was entirely self-consistent; it was just consistently wrong about an obligation that had already been discharged. Self-consistency cannot see that. Only an independent answer can.

That is the argument for the corpus, and it is also the reason not to trust the property tests alone.


2. Nobody has used this

No customer has run this tool on a real project. Every claim about what reviewers want comes from documents — specifications, standards, published research and one InTrans/FHWA study of eleven state DOTs — not from reviewers.

~~There is no user interface.~~ Corrected 6 September 2026. There is a Python library, a command-line entry point, and construct serve — a local page on 127.0.0.1 over that same command line, which computes nothing of its own. What there is not is anything hosted: nobody can be sent a link, so trying it means installing it first. outreach/OBJECTIONS.md §9 states it in that form.


3. ProgressMode.ACTUAL_DATES is an interpretation, not a reproduction

Primavera offers three out-of-sequence progress modes. Two of them — retained logic and progress override — have well-understood behaviour. The third does not, and the rule implemented here:

suppress a logic tie where the recorded dates already contradict it: the successor has an actual start, and the predecessor either has no actual finish or finished after that start

is this project's reading of "drives from recorded actuals rather than forecasts", chosen because it is the reading that makes the mode distinct from the two either side of it and produces the strict ordering finish(override) ≤ finish(actual_dates) ≤ finish(retained).

It has not been validated against real Primavera output. The reasoning is recorded in engine/tests/test_progress_modes.py rather than in a commit message, so that whoever runs the P6 validation knows exactly which claim to check.

Two consequences for a reviewer:


4. Concurrency returns no governing answer under US jurisdictions

Under us_federal, us_state_or_private and neutral, the concurrency module's governing() returns None. It reports all four doctrinal findings, states which authority declined and why, and names the open question.

That is not a missing feature. AACE RP 29R-03 §4.2.D.1 sets out the literal and functional theories without endorsing either, and §4.2.D treats the choice as one for the analyst to make and defend. No document in the US precedence supplies a rule that selects between them. A tool that picked one to avoid returning None would be manufacturing the certainty that gets an expert excluded rather than merely disagreed with.

Only the SCL Protocol has an encoded position, so only uk_scl produces a governing finding — the functional / first-in-time reading, per ¶10.4 and ¶10.10.

Related limits in the same module:


5. Source Validation Protocol gaps


6. Conformance coverage


7. Methods not implemented

All nine MIPs now run, plus prospective TIA under RP 52R-06.

That is a statement about implementation, not about correctness. "Available" means the engine performs the method end to end and there are tests; it does not mean any result has been checked against Primavera, against a published expert analysis, or against a real project. Five of the nine were built in a single session and none has been run on a real schedule.

The limitation that matters is therefore not which methods exist but §1 of this document: the engine is not validated at scale. A method that runs on generated fixtures and has never met a 4,000-activity update with real progress, real calendar exceptions and real out-of-sequence work is a method whose failure modes are unknown.

Read forensic.methods.catalogue() rather than this paragraph — it reports the same table from the code and cannot go stale the way prose does. See METHODS.md.


8. Input formats and scheduling settings


9. Where the code and its own documentation disagreed

Eight disagreements were found while writing this documentation. All eight have since been resolved, and the record is kept because how they arose is more useful than the fact that they did: every one was a place where a string, a comment or a default asserted something the code did not do, and none of them would have been caught by a test of behaviour.

  1. collapsed_as_built labelled itself MIP 3.9 and performed MIP 3.8. Fixed. One network, one set of events, one collapse is modelled / subtractive / single simulation, which is 3.8; 3.9 is the period-by-period collapse. The label survived because a test asserted it — assert "3.9" in result.mip — so every run confirmed the error. tests/test_method_labels_match_the_taxonomy.py now checks the engine's self-labels against forensic.mip, in the forensic layer because the import contract forbids cpmcore from seeing the taxonomy.

  2. forensic.svp.REQUIRED_TERMS had drifted in both directions. Fixed. contract_scope and contract_terms were read and undeclared, so a reviewer had no way to know they were wanted. update_submittal_log was declared and read by nothing, which is worse: a reviewer supplies it, nothing happens, and concludes the tool ignored their evidence. It is removed, and the check that would use it — distinguishing submitted updates from working copies, SVP 2.3.B — is recorded as unbuilt in §3 rather than implied by a table of inputs.

  3. The command line asserted materiality it never measured. Fixed. It marked every disclosure material, so the report printed "3 of 3 choices below would change the result on this schedule" — a sentence asserting something no computation established, in the reporting layer, which is precisely the failure this package exists to prevent. Disclosure.material is now tri-state; None means nobody established it, and the renderer has its own section saying so in those words.

Where it was actually fixed, because the first entry did not say and that is how the review found it still open: Disclosure.material now defaults to None in forensic/disclosure.py. The earlier fix changed only the CLI's call site, so every other caller still asserted materiality by omission. An entry that reads as closed stops the next reader checking, which makes a half-fix more dangerous than none.

  1. The conformance section's confidence label was tied to coverage. Fixed. COMPUTED_UNVALIDATED means "computed from inputs that could not be validated", which is true of every run of this command regardless of coverage, because the command does not perform the Source Validation Protocols. Coverage measures something else entirely.

  2. UFGS clause 073 was unaccounted for. Fixed. It is a second enforcement point for §3.12(f), which runs as UFGS-098; running both would report one defect twice under two clause numbers, which reads to a reviewer as two problems. The pack's header now says so.

  3. engine/oracle/README.md was stale. Fixed. It now reports 156 of 160 and four named divergences, matching the harness.

  4. forensic.rules.asce67 exported only REGISTRY. Fixed. ~~It now exports RULE_COUNT, registry() and counts_by_determinism() with the same shapes as the UFGS pack — two packs behind one interface should not need two call shapes to answer the same question.~~

Superseded 6 September 2026. The struck sentence is kept because it is the record of what the fix was. Two of the three names it publishes no longer exist: registry() and counts_by_determinism() were deleted from all thirteen packs, 26 functions, having no caller anywhere under engine/src, engine/oracle, engine/quality or demo. A pack's registry is reached as the module constant REGISTRY, which is what forensic/cli.py builds its authority packs from and what forensic/rules/dependence.py discovers packs by; a pack's size is RULE_COUNT. Those two names are the pack interface, they are asserted for every pack by engine/tests/test_pack_interface.py, and the original point stands unchanged: thirteen packs behind one interface do not need thirteen call shapes to answer the same question. The per-determinism census that counts_by_determinism() returned is Counter(r.determinism for r in PACK.REGISTRY.rules) at any caller that wants it. Recorded in engine/quality/UNREACHABLE.md, section 7.

  1. forensic.mip.candidates() and all_profiles() were public and missing from __all__. Fixed.

10. What none of this is

Nothing produced by this tool is legal advice, and nothing in it is advice about any project. The concurrency and jurisdiction modules encode what published documents say, with the paragraph references so a reader can check, and leave the places they disagree disagreeing.

The FRE 902(13)/(14) certification the evidence module produces is draft text for a qualified person to review, adopt and sign. It is not a certification until a person with knowledge signs it, and this engine cannot be that person. Whether any given signatory is "qualified" within the Rule is a judgement for counsel. Self-authentication addresses authenticity only, not hearsay — a self-authenticated record still has to get past FRE 802 — and Rule 902(11)'s procedure, which (13) and (14) incorporate, requires reasonable written notice to the opponent and an opportunity to inspect. Missing that step forfeits the benefit.

Gap detection states observable facts about the set of documents provided and nothing more. It does not assert that anything was lost, destroyed, withheld or concealed, because FRCP 37(e) makes that a question of preservation duty, reasonable steps and — for the severe sanctions of 37(e)(2) — intent to deprive, none of which is visible in a set of files. The most common innocent explanation for a missing update is that nobody sent it.

11. What the delay-injection matrix demonstrated a method cannot do

Added 2026-09-02. Every other oracle tier in this repository measures the scheduler against P6 or the reader against MPXJ. engine/oracle/injection/ measures the forensic layer directly: it injects a delay of hand-counted size into a five-activity network and asks whether each MIP reports the delay that was injected. engine/oracle/injection/MATRIX.md carries the run, the predictions registered before it, and the mutations that verify it can fail.

All five exercised methods recovered every headline figure, including a placebo (nothing injected, every method reported zero) and a neutral control (a float activity extended five working days against fifteen days of float, every method reported zero). The limitations below are what the same run showed about attribution, and each has a demonstration behind it rather than a reservation.

MIP 3.3 stops explaining a delay at the point the delay becomes decisive. Two seven-day slips on the same fixture: where a critical activity was extended, the windows analysis attributed every day to a named cause on a named activity; where a float activity overran far enough to take over the driving path, it attributed nothing — the whole movement landed in the residual PATH_SWITCH / UNEXPLAINED. The magnitude is right in both. So explained_fraction reading 0.0 can mean the analysis did no work and can mean the delay was large enough to change the critical path, and this method does not distinguish them. A report quoting an unattributed figure has to say which the reader is looking at, and no method here can tell them.

MIP 3.1's per-activity table counts float consumption as a delayed finish. On the neutral control the project figure is correctly zero and candidate_as_built_critical_path is correctly empty — exceeds_planned_float is doing the RP's late-date test, as engine/quality/MIP-COVERAGE.md §2 records. But delayed_finish and extended_duration are both True for the float activity, so a reader summing delayed_activities obtains five days of delay on a project that was not delayed. The headline is right; the table beneath it is a list of movements, not of delays, and is not labelled as one.

MIP 3.2's per-period variance was not the delay that accrued in that period. Fixed 2026-09-02; the reservation is withdrawn. On a seven-day delay the two periods published +21 and −14. They summed to the gross figure and RP §3.2's reconciliation held; what a reader was handed was twenty-one days of variance in a period ending before any of it happened. Activities were filed into a period by their as-planned finish while the level was taken over their as-built finishes, so any float activity overrunning its period contributed the whole overrun to the earlier period and a negative correction to the later one — and Period.variance_days reads a negative as time recovered, which on that fixture nothing was.

An activity is now filed on the later of its two finishes, so a period's figure can only draw on dates its own boundary has passed. The same injection publishes [0, +7]. Filing on the as-built finish alone was rejected: it has the same defect with the signs exchanged — an activity planned late and finished early would file into an early period and drag its planned date forward out of it, claiming a recovery weeks before the work it was recovered against was due. A genuine recovery remains representable and is tested for (engine/tests/test_period_filing.py), which reads the per-period figures: MATRIX.md §6 records that a check reading only the total cannot see a filing defect. MATRIX.md §4.2 still describes this as an unfixed defect candidate; that section is the record of the finding and is superseded here.

A constraint-driven delay has no MIP 3.6 input form at all. cpmcore.modelled.DelayEvent attaches to the network as a Finish-Start predecessor and carries no constraint field, so a delay imposed by a date constraint — ten working days on the fixture, recovered in full by all four observational methods — cannot be handed to the impacted-as-planned method in any form. That is n/a rather than a miss, and it bounds what a modelled analysis can be asked about on a schedule whose delays live in its constraints.

What the matrix does not establish is in MATRIX.md §8 and is short enough to repeat: five activities, one injection at a time, no concurrency, no fuzz, and no attribution of cause to a party anywhere — the engine publishes attributes_cause: false and means it. MIP 3.5 was not exercised; §7 there says why.

Superseded on 2 September 2026. That paragraph also said 3.7, 3.8 and 3.9 were not exercised, and that the matrix had no calendar mix. Both were true when written and are no longer: the collapse methods were injected into the first fixture in this repository that satisfies §3.8.E.1, on a network carrying a five-day and a seven-day calendar, and engine/oracle/injection/MATRIX-COLLAPSE.md has the runs. The four bounds below come from it.

11a. The collapsed as-built — four bounds, each demonstrated

The total depends on an anchor nobody states. Both sides of a collapse are floored at a data date. MIP 3.8 infers one from the earliest actual start; MIP 3.9 cumulative uses the first base's. On the reference fixture the same model and the same three delay events give fourteen working days at one anchor and eleven at the other, and the concurrency figure moves with it — plus five against plus three on a second run. Ask which anchor a collapse figure was produced at. SubtractiveMultiBaseResult.anchor_note states it in the record.

A small delay near a calendar boundary can be recovered as zero, correctly. One working day of delay on a five-day-calendar activity feeding a seven-day one is recovered as one day at the inferred anchor and zero at the hard anchor — because the project finish genuinely did not move; the day landed inside a weekend the successor was waiting through anyway. Do not read a zero as the method having missed something. It also means a dose curve for MIP 3.8 is a property of the anchor as much as of the method.

A delay recorded in line, with no §3.8.K.2.c redundant tie around it, produces a collapse larger than the delay removed. Thirty-three and thirty-four working days out of eighteen days extracted, with a successor finishing before its own predecessor — and §3.8.E.1 still reads PASS. The only disclosure is the severed-chain table, which names every broken pair and is correct. Read that table before the total: if it is not empty, the total is not a delay figure.

A per-event figure is not an apportionment. Lengthening one delay changes another's reported contribution without that second delay being touched — on the fixture, one event reads four days and then three. A single-event collapse is a statement about that event in that programme, not a property of the event, and the per-event figures do not sum to anything a reader should quote.

The bound on all four. These come from fixtures built to satisfy §3.8.E.1, with their dates copied out of the same scheduler then asked to reproduce them. A PASS is evidence about the method, never about the population. Of the sixty-nine corpus files, sixty-six schedule; of those, forty-eight record no actual dates at all and eighteen record them and fail. None passes.

One thing the collapse family cannot have, and it is a fact about testability rather than a defect. There is no clean placebo on a mixed-calendar as-built. Dosing every delay to zero and extracting them still moves the finish by one day, because a recorded zero-duration activity is still a calendar boundary — its seven-day successor waits for the Monday its five-day calendar imposes. The finish really did move, so the method is right. It means the cheapest control in the whole design is unavailable here, and the placebo that is available has to be read with that in mind.

11b. The method is an input to the answer, and here is what that costs

Every honest account of forensic schedule analysis says the choice of method affects the result. It is usually said as a caution. It can now be said as a number, on one fixture, on one day, with both figures hand-counted.

The same delay event is worth nothing and worth a week, depending only on which method is asked. DLY-APPROVAL — an owner-caused approval delay that was real, was recorded, and never drove the project finish — is worth:

Neither figure is wrong and neither method is misapplied. They answer different questions: what would have happened if this had not occurred against what happened when this was introduced. On a schedule where the driving path changes hands — which is most real schedules — those questions have different answers, and the difference is not small. Here it is the whole of the claim.

What follows for anyone quoting a delay figure from this tool. A number without its method is not an answer. The engine publishes the method with every figure and attributes_cause: false alongside it, and neither of those is boilerplate: this is the case that shows what they are protecting against. If an expert on the other side runs a different MIP and gets a different number, that is not necessarily either party being wrong, and a report that does not say so is inviting a cross-examination it cannot survive.

Two more from the same run, recorded because they bound what MIP 3.7 can be asked. Its dose curve terminates — at 22 event-days on this fixture, where one event's duration grows past the next event's onset, §3.7.E.7 fires and the method withholds the total rather than publishing one. That is a correct refusal. And the double-count risk of impacting an event into a base that already contains it does not reach any figure 3.7 publishes: both sides of the subtraction move together, so combined_days and the total hold at 6 and 14. Only the finish moves, 2026-05-12 to 2026-05-20 — and §3.7.E.14 already declares that finish is a figure the method cannot check.

12. Output surfaces that are still hard to read aloud

Audited whole on 2 September 2026, for the first time (docs/ACCESSIBILITY.md, which owns every finding and is the page to read; this entry exists so that "everything the tool cannot do" does not silently omit the one axis its own author reads on). The Markdown report, the HTML report and the JSON record came out sound. What did not, and is recorded there rather than here because it moves: two of the eight error and refusal paths do not read as sentences, and one warning in forensic.evidence verify dumps a Python repr at the reader. That last one is a stated tension and not an oversight — accessibility and stating evidence precisely pull against each other there, and it was left visible rather than resolved by quietly dropping the precise form.

This matters commercially and not only ethically: a forensic report is read aloud, quoted in a letter and pasted into a filing, so text that works only visually is broken output rather than ugly output.


13. Every check that needs a second machine has never run

Recorded overnight 5-6 September 2026. This entry is here because the rest of the document describes checks that execute, and the checks in this section do not. A check that has never run is not weaker evidence than a check that has. It is not evidence.

Continuous integration has been dead since 3 September and 284 commits have landed since. .github/workflows/ci.yml defines five jobs — check, differential, package, record and reproduce — and it is the half of the verification that needs real interpreters, real runners or a network, which is to say the half nobody would type by hand. GitHub reports 429 workflow runs for this repository. Of the 200 most recent, none concluded success; the oldest of those 200 was created at 2026-09-03T09:04:05Z, so the failing streak covers every run in more than two days. Runs on today's commits conclude failure in four to seven seconds, which is the shape of a job that never started a step. engine/quality/CI-AUDIT.md §0a has the cause in GitHub's own words: "the job was not started because recent account payments have failed or your spending limit needs to be increased." It is a billing state and it is the one defect in this repository that cannot be fixed from inside the tree.

The reproduction matrix has never executed once, in the life of the repository. reproduce is declared needs: record, and record has never passed, so the matrix has been skipped on every run of its existence (CI-AUDIT.md §1, which measured that before the billing halt and again after it). That matrix is five legs — Ubuntu and Windows against Python 3.11, 3.12 and 3.13, less the leg that writes the record — and it is the only place two legs of this project's central claim are measured at all:

The three-way differential against MPXJ has never run automatically either (CI-AUDIT.md §1). It is the only check in this project that can say we are wrong without asking Primavera, and it has only ever run when somebody typed it.

There is no install route but a clone. pip install construct-engine and pipx install construct-engine resolve to nothing; the name is unregistered, and docs/RELEASING.md §9 records the date and the HTTP status of that check. The package job — which builds the wheel and installs it the way a stranger would, with --no-index — passed exactly once, on 3 September, before the halt. docs/INSTALL.md states the clone requirement on its first screen and is the page to read; it is repeated here because "what this cannot do" should not need a reader to have already found the installation page.

What this does not license anyone to say. Nothing above means the engine is non-deterministic, or broken on 3.12, or unusable on Linux. It means nobody knows, and the difference between unmeasured and fine is the whole subject of this document. The gate that a developer can run — engine/check.py, the unit suite, the oracle suite, the import contracts and the P6 agreement — is unaffected by all of this and runs on every machine that has a clone.

What this tool may not be used to produce

A limit that is not ours and cannot be engineered away. UFGS 01 32 17.00 20 (NAVFAC, 05/25) makes third-party processing of an XER cause for rejection. The submission a contractor sends must come from their own Primavera, not from anything this tool wrote.

That forecloses a whole class of feature permanently, on exactly the federal work this product targets: no corrected file, no re-export, no "fixed" schedule handed back, however carefully any of it were built. cpmcore ships three writers (write_xer, write_p6xml, write_mspdi) and no shipped command reaches any of them -- an accident of history that turns out to match the constraint, and which should stay that way unless someone can name a use outside the differential oracle that this clause permits.

It also bounds the honest shape of anything contractor-facing. respond.py groups a report's findings by what the recipient must do and names the missing inputs an abstention depends on; it produces prose, never a file, and docs/CONTRACTOR-MODE.md records why that is the ceiling rather than the first version.

Recorded here 2026-09-02. The clause was established in corpus/analysis/FIELD-EVIDENCE.md and cited in business/ and docs/CONTRACTOR-MODE.md, and was absent from the page that owns limitations -- found by an agent that grepped this file for it while designing against it.

Source: docs/LIMITATIONS.md. Source commit date: 2026-09-06.

See it in practice

Follow the evidence, from the schedule to the finding.

Explore the worked example