Limitations
Everything this tool cannot do, or has not proved. Consolidated here so that no user is surprised, and so that this page can be pointed at.
A company selling evidentiary rigour that overstates its own validation has destroyed the only thing it sells, and would be found out by exactly the audience it was trying to impress. So the position is stated first and in full.
1. The engine is not validated at scale
Updated 1 September 2026. This section previously said no held-out corpus
paired real schedules with P6's own computed answers, and that the 27-activity
fixture below was the only real P6 evidence held. Both statements were true when
written and are now false: engine/oracle/corpus/ holds eleven harvested P6
exports, and agreement across them has been measured. The measurement is worse
than the fixture's, which is why the correction is here rather than in a
changelog.
The ceiling on this measurement has itself been measured, and it does not
excuse us. P6 was run against P6 over the same 584,687 comparisons with our
scheduler never called: 11,330 of them sit on a row whose stored values cannot
all be true at once, so achievable agreement on this corpus is capped at 98.06%
(engine/oracle/corpus/P6-SELF.md). Of the 32.14 points short of 100%, at most
1.73 are P6's and at least 30.41 are ours — the apportionment was tightened
against us on 3 September, when joining the two instruments per cell showed that
1,218 of the 11,330 unreachable cells are ones we already agree with and so are
not recoverable ground. That is the sentence to read the
rest of this section through — a reader who sees the level below without it
cannot tell an engine carrying thirty points of defects from one measured
against a noisy target, and the answer is that the target is not noisy. The
self-inconsistency does not even fall where our gap falls, and every
self-inconsistent project already fails the validity gate.
Agreement with P6 is 67.85% — 396,735 of 584,687 field comparisons across 82
projects and 86,031 activities, P6 5.0 to 24.12
(engine/oracle/corpus/SWEEP-2.md, which measures and owns this figure; this
page cites it rather than asserting the level as its own). The per-project table
in that document is the honest presentation. Over the 66 projects that pass the
validity gate the
figure is 78.06%, and that belongs beside the headline rather than instead of
it: the 16 projects the gate refuses hold 40.9% of all comparisons, and a harness
that improves its number by choosing what to measure has stopped measuring. The
gate was itself measured on 1 September: its duration-against-span check had
reached one corpus project's calendar on 0 of its 770 unstarted rows and recorded a
PASS anyway, because the harness read the week as a union of shift shapes and
summed overlapping ones into a fifteen-hour day on a ten-hour calendar. Reading
the per-weekday shifts instead takes it to 770 of 770; 494 of the 507 misfits
were the instrument, 14 are real, and the file moves from an unearned PASS to
a FAIL. It leaves the gate-passing population, taking it from 65 to 64,
and because it agreed on only 40.5% of its comparisons that raises the
gate-passing figure by 0.63 points — arithmetic, not improvement. The
falsification was run before the percentage precisely because the movement
flatters us: 494 of the 507 union misfits were overshoots by exact multiples of
2.5 hours, the 14 survivors are undershoots of one to eight hours in one
contiguous sub-network, and a residual matching the same signature would have
meant the fix was tuning.
This paragraph said 493 and "all 507" until 2026-09-02, and the arithmetic
that looks like it justifies 493 is a trap. 507 − 14 = 493 assumes the 14
residuals are a subset of the 507 union misfits. They are not.
engine/oracle/harness.py:1072-1080 states the split: 13 of the 14 already
missed under the union, by the amounts that were not multiples of 2.5, and the
fourteenth was scored as reconciling exactly by the union model and does not
reconcile — a right answer for no reason at the scale of a single row, and the
reason a corrected instrument is allowed to find more faults than the one it
replaced. So the 507 partition as 494 instrument plus 13 real, and the residual
is 13 + 1 = 14. engine/oracle/corpus/SWEEP-2.md:186-192 had it right and this
page had it wrong, which is the reverse of the usual direction here. The gate now reports PASS 64, UNPROVEN 2, FAIL 16 — it read PASS 62 / UNPROVEN 2 / FAIL 18
until the base-calendar fix of 2 September took two files from FAIL to PASS — where
UNPROVEN is a strict subset of the 66 passing rather than a third population
(engine/oracle/corpus/SWEEP-2.md, engine/oracle/harness.py). The
scheduling logic is internally consistent, deterministic and property-checked,
and that is not the same thing as agreeing with the source software. Figures
should be reconciled against P6 before they are relied on.
By the end of 1 September 2026 it had moved three times, in opposite directions and for opposite reasons, and they must not be run together; four further movements followed on 2 September and are set out below. It was 79.32% on a population of 22 projects and fell to 60.51% when the corpus grew fifteenfold: that fall is the correct result and not a regression, because the old population contained not one mid-project update and so measured the engine only on the easiest schedules P6 produces. It then rose to 61.77% on that same fixed corpus when a real arithmetic defect was fixed — a start-to-start lag charged twice where the predecessor had already started, which P6 has already spent against the actual start. Comparison, project and activity counts are identical either side of it; 7,392 comparisons came right.
It rose again, from the superseded 61.77% to 61.89%, on a second fix at
9188ad8. xer.py converted every duration through a calendar's declared
day_hr_cnt, and eight CALENDAR rows in five corpus files contradict it — an
eight-hour day on a calendar running 07:00 to 17:00. Five of the eight carry
activities and those five are four distinct calendars, one appearing in two
projects; the other three carry none. Every duration on the affected calendars
was stretched by a quarter.
Two separate engine defects moved this figure on 1 September 2026, in two
movements: 60.51% to 61.77%, then 61.77% to 61.89%. Three figures are two
movements, and this page said "three defects" for a day because a third commit
of the same date — f5b4fb7, which floors unstarted work at a project planned
start later than the data date — was read into the chain. Its own message states
what it moved: the eleven-file corpus, 55.11% to 79.32%, before the
fifteenfold expansion, with the cross-validation fixture unmoved at 156/160. It
never touched the 82-project figure, and SWEEP-2.md names the two that did.
The reading does not change with either of them. Defects of that size
sitting in a lag rule and a calendar's declared day length say how much remains,
not how far this has come, and on that day 38.11% of 584,687 comparisons still
disagreed. All three of 60.51%, 61.77% and 61.89% are superseded as headlines
and must not be quoted as this engine's agreement; the 79.32% in that commit
message is superseded too, and was never a figure for this population at all.
It rose again the next day, from 61.89% to 67.50%, and that movement was
larger than every scheduler fix before it put together. Neither defect was in
the scheduler. _read_calendars never followed CALENDAR.base_clndr_id, so a
project calendar did not inherit its base calendar's exception days — worth
66.73% alone, moving exactly two projects, one of which
went from 10,882 to 38,904 of 39,256 and
from 1,304 misfits on its duration-against-span check to 0. And P6's
midnight act_end_date is a day start, not a day end — worth 62.66% alone,
moving exactly the seven projects WHOLEDAY.md had named in advance. Both are
fields P6 stores; neither is a P6 answer read back in. Every per-project
prediction registered before the fixes held to the digit.
Read it with the number that fell. Importer fidelity arm 1 — agreement with
MPXJ, an independent importer — went from 99.60% to 99.43% over the same
commit, and it reconciles exactly: 1,408 rows of a new ACTUAL_FINISH class
where MPXJ keeps P6's raw midnight stamp and we now read it as a day start, plus
45 inherited exception days MPXJ does not follow. Arm 1 measures agreement with
another importer's convention, not with P6's stored dates, so it is not a
measure of correctness and a fall in it beside a rise against P6's own
arithmetic is the expected shape (engine/oracle/FIDELITY.md).
It moved twice more the same day, and only one of those two is agreement.
Every level named in this paragraph is superseded and is here as history:
these are the movements, not the current pair, which is at the top of this page.
422cc25 took it to 67.54% by reading the second of P6's two constraint slots
in the forward pass — +243 comparisons on one project, arithmetic. The move that
took the headline from 67.54% to 67.71% is not that, and the matching move of
77.60% to
77.85% over the projects that pass the gate is not either — and both of those
levels are superseded, the gate-passing one having since gone 77.85% to 78.06%
in three further steps set out below. P6's
driving_path_flag marks membership of the driving set — the backward closure
over driving relationships, which branches at every tie — and is_longest_path
had been marking membership of a chain, one predecessor per hop. A chain is a
subset of the set it lives in, so the disagreement could only ever run one way,
and it did: 2,564 rows to 94. Publishing the set drops the DRIVING_PATH
disagreement class from 2,658 to 1,628 while every other class is unchanged
count for count. The same evidence, scored against a corrected reading of the
question — not more agreement on the same one, and this page says so because a
sentence reading "agreement rose" would describe a measurement that did not
happen.
Both figures moved once more later the same day, 67.71% to 67.80% and
77.85% to 77.97%, for the reason the next paragraph gives — and once more again
on 2026-09-04, 67.80% to 67.86% and 77.97% to 78.07%, on a start-to-finish lag advanced in start space and read back as a finish boundary, which costs a working day whenever the day before the advanced instant is not worked. That
fifth movement is the first that is arithmetic in the scheduling kernel; P6 and
MPXJ both give the answer we now give on the activity that names it. And once more on 2026-09-06, and it is the first that goes down: 67.86% to 67.85% all-projects and 78.07% to 78.06% over the gate-passing projects, when a finish-side constraint date was read as the boundary it names rather than as the start instant of its day. 27 comparisons in two projects, both falling, with the other 80 holding cell for cell. It was kept and reported falling because BOUNDARY-SEMANTICS.md settles the reading from P6's own stored bytes.
A fourth movement followed, later the same day, and it is a fourth kind of
thing again. c0c69ef had been reading a stored constraint stamp with a start
rule where the constraint needed a finish rule — the same instant, resolved to
the wrong end of the working day. Reading it on the correct day took the headline
from 67.71% to 67.80%, which is 526 comparisons across all 82 projects. Over the
66 projects that pass the validity gate the same fix moved 77.85% to 77.97%, on
409 comparisons — a level since superseded in turn, so that the chain over those
projects reads 77.85% to 78.06% across the three movements since. Every one of the 526 lands in a class carrying a mid-day stamp or none
at all, and all three whole-day classes are unchanged to the row — which is the
check on the mechanism rather than on the number. One project fell by exactly
one comparison and was kept, two of its rows having moved onto P6's own stored
answer; a correct fix that lowers a number is reported falling, not reversed.
Four movements on 2 September 2026, and no two of them the same kind of thing: two importer defects, then arithmetic, then a change of definition, then a stamp read on the wrong day. Six movements in two days, counting the two of 1 September above. A page presenting these as one number improving six times is telling a false story out of correct digits, and no check in this repository can catch that — which is why the chain is written out here in sentences rather than left as a row of figures.
A single arithmetic defect was worth 1.26 points. That is not a milestone.
It says how much is still wrong, not how far this has come, and close to a third
of 584,687 comparisons still disagree. Two figures keep it in proportion: the
SWEEP-2.md sweep-population rows, 80.0% and 79.32% — superseded as headlines,
still correct for their own narrower populations — did not move at all,
because none of those files carries the feature — an independent check that the
fix touched nine projects and not the corpus; and the per-project median went
85.7% to 85.6%, falling while the aggregate rose, because a median does not
feel a fix that lands in nine projects out of 82.
A named architectural cause is on record and is not a bug: the time axis is
whole-day, and an in-progress activity with a fractional remaining duration
finishes at a mid-day instant in P6 that this engine advances to the next whole
working day. engine/oracle/corpus/APPORTIONMENT.md measures that at 31.3% of
the gate-passing gap (23,830 of 76,081), from its class table re-run on
2026-09-02 over the 66 that now pass; that document's two SUBDAY sub-analyses
are marked pre-fix and are not quoted here.
Two things about that figure, and the second is a coincidence a reader will otherwise misread.
First, ~~30.9% (23,877 of 77,376)~~ is superseded. It was correct against the gap as it stood before the constraint clock landed, and the owning document has since been re-run against a frozen tree -- reproducing all eighteen of its previously published class counts cell for cell before publishing any new one.
Second, this paragraph read 31.3% once before, and that was a different
measurement. It moved to 30.9% when the class table's own instrument was
corrected earlier the same day: gapshape.day_bounds collapsed a calendar's
whole definition to a single boundary pair, so a stamp that is a day end on its
own day block read as mid-day. 4,202 rows changed column and none changed
verdict, which is why that figure moved while no agreement figure did. The
correction has not been reverted -- the number is back at 31.3% because the
denominator fell from 77,376 to 76,081, not because the classifier changed
back. Two different measurements agreeing to a decimal place is the kind of
coincidence that makes a page look stable while it is being rewritten
underneath, which is why it is spelled out rather than left as a digit.
And the count fell while the share rose. Sub-day lost 47 rows, 23,877 to 23,830. Its share went up because the gap fell faster. Any sentence reading the rising share as this class growing describes a measurement nobody made -- and would pass every digit-comparing check in this repository.
The shares moved with the importer fixes and the ranking moved with them.
WHOLEDAY_BEYOND_ONE_DAY went from 44,915 rows to 23,551, the gap fell from
222,803 to 190,046, and the population passing the gate went from 64 to 66 — so
every share on this page is a share of a smaller gap over a different
population. Aggregated by family, the mid-day classes went from 44.9% to 51.9%
of the gap and the whole-day classes from 27.6% to 19.3%: the fix that raised
the headline made the mid-day family a larger share of what is left.
That class table is an 08cc4e4 measurement and has not been re-run since
the definition change or the constraint clock. Its DRIVING_PATH row of 2,658
is now 1,628, and the constraint clock then moved four more rows — SUBDAY by
141, MIDDAY_BEYOND_ONE_DAY by 101, TOTAL_FLOAT by 247 and FREE_FLOAT by 37
— so five of the nine counts have moved since 08cc4e4, and the 190,046
total is now 188,247. Read that span carefully: across the constraint clock
alone it is four classes that moved and five that held, and it is only
cumulatively from 08cc4e4 that five of nine have. SWEEP-2.md puts sub-day's
share of the gate-passing gap at 31.3% on the live figures — 23,830 rows of
76,081, where this page previously read 23,877 of 76,490. The share rose while
sub-day's own count fell by 47 rows, because the gap it is a share of fell
faster. That is the arithmetic of the denominator, not sub-day growing, and a
sentence reading it as growth would pass any check that compares digits.
A class here can grow while the gap shrinks, and SUBDAY has now done it
twice — 39,614 rows to 44,645 on 1 September, and 44,645 to 45,230 on
2 September while the gap fell by 32,757. A row corrected from many working days
out lands one day out rather than vanishing, so it moves into the sub-day
class. That reads as a regression to anyone not told why, and it is not one.
No class is largest by a margin worth publishing, and this page will not say
one is. Across all 82 projects the top three are SUBDAY 23.7%,
TOTAL_FLOAT 23.4% and MIDDAY_BEYOND_ONE_DAY 22.0% — 44,526, 44,016 and
41,457 rows in a gap of 188,247. The first two are 510 comparisons apart,
which is less than the 526 a single fix moved that same day; the third sits
3,069 behind the first. The figure this page carried before, "a spread of 310",
was the spread of the top two on the superseded counts and was never the
spread of the three named beside it.
APPORTIONMENT.md puts it as a tie and says in terms that any sentence of the
form "X is the largest class" is a claim about a coin flip. The scale that makes
that true: a single importer fix moved WHOLEDAY_BEYOND_ONE_DAY by 21,364
rows, and a correction to the classifier -- not to the engine -- moved 4,202
and swapped the second and third places recorded here. Over the 66 projects
that pass, SUBDAY holds 31.3% of the gap (~~30.9%~~ superseded by the
2026-09-02 re-run of the owning document).
The second place has now swapped again, and it is worth saying how narrowly.
TOTAL_FLOAT is at 24.5% and MIDDAY_BEYOND_ONE_DAY at 24.4% — over
the gate-passing population that is 46 rows apart in a gap of 76,081, a
tenth of a point. This page previously paired SUBDAY with
MIDDAY_BEYOND_ONE_DAY as the top two; both of those digits were individually
correct and the pairing had already gone stale. Two classes that have
reordered twice in one day, on changes that moved no verdict, are not a first
and a second. Do not quote a second-largest class from this page, and treat
any sentence naming one as a claim about a coin flip — which is what the owning
document says in terms.
The earlier version of this paragraph expressed that scale as a multiple of
the spread. It is stated in rows now: the spread it divided by has itself moved
twice, and recomputing the multiple would publish an arithmetic nobody measured
-- which is the failure this page exists to describe, committed on the page
describing it. Earlier versions of this paragraph named a largest class three times
and were wrong twice. MIDDAY_BEYOND_ONE_DAY's mechanism
is established as of 1 September 2026, which it was not when this paragraph
was written.
The mid-day handoff. When P6 finishes a predecessor part-way through a day it starts the finish-to-start successor at that same instant on that same calendar date — both activities run on 16 April. A whole-day engine cannot: if the predecessor occupies the 17th, the successor starts the 18th. The successor therefore loses a second working day at the link, on top of the first lost within the activity by rounding its finish up.
That result is controlled, which is what makes it causal rather than
correlational. Taking only activities all of whose predecessors' finishes
already agree with P6, so nothing is inherited: 5,441 activities in 22 files
hand off at a mid-day instant, and the successor's start is one working day late
on 5,429 of them — 99.8%. The matched control is the identical topology at a
shift end: 1,974 activities in 27 files hand off at a day boundary, and 1,944
of them — 98.5% — agree exactly. The timestamp is doing the work, not the
shape of the network (engine/oracle/corpus/MIDDAY-MECHANISM.md).
So the "sub-day costs at most one working day per activity" bound is false. The replacement is bounded rather than open-ended, and it is measured rather than controlled — a carry model over 50,707 disagreeing unstarted rows, not the control population's own result:
A whole-day time axis costs at most one working day per activity in 97.8% of cases and at most two in 99.4%, and the second day arises at the link rather than within the activity.
Read the two weights differently. The 5,429-against-1,944 result is observed with a control. The 97.8 / 99.4 figures are modelled over the corpus. And the residual — 306 activities, 0.6%, where the model predicts three working days and the median is fifteen — is explicitly not claimed: the carry model reaches those rows without explaining them.
This is not a defect awaiting a patch. A pass that set out to write the fix
measured the input it needs and refused it (MIDDAY-MECHANISM.md, "The bit
cannot be derived", reproducible by python derive.py):
This share of the gap is a property of the whole-day time axis rather than a defect awaiting a patch: the quantity that drives it is sub-day, it is inherited across activities rather than generated at them, and reconstructing it from whole-day inputs recovers at most 70% of the affected rows while misfiring on 14% of the rows currently correct.
Each clause is measured. Inherited, not generated: on the document's own named case the offset arrives across eleven hops whose durations are every one a whole multiple of the eight-hour day — 144, 96, 48, 24, 24, 56, 56, 56, 120, 32, 24 — from an in-progress activity with 11.2 remaining hours; across the file, only 1,093 of 8,962 mid-day handoffs — 12.2% — have a fractionally-durated immediate predecessor, so seven of every eight inherit it from further upstream. At most 70%: a derivation built deliberately stronger than any engine could be — carrying the offset exactly in minutes, and choosing the driving predecessor using P6's own stored dates, an advantage this engine does not have — fires on 3,807 of the 5,441 mid-day rows, 70.0%. Misfiring on 14%: the same derivation fires on 280 of the 1,974 day-boundary control rows, 14.2%, breaking about 280 of the 1,944 starts we currently get right.
That defeats the fix as specified. It is not a proof that no derivation exists, and the document does not claim one.
None of this is good news, and it should not be read as any. No agreement figure moved for this reason: the mid-day work changed no behaviour, and the corpus documents say in terms that nothing in it disturbs the apportionment. The headline did move afterwards, from the now-superseded 60.51% to 61.89%, and it moved because of an unrelated start-to-start lag fix — not because anything here was applied. Two things moved in opposite directions in the same pass: a hedge became a bounded statement with a control behind it, and the prospect of fixing it receded from "specified, pending" to "defeated as specified". Nor does the mechanism explain the class — even in the file where it was found, the carry accounts for the early-start error on only 1,758 of 11,368 unstarted activities, 15.5%, and the largest single unexplained thing in the corpus is now 230 forward-pass origins in that file whose predecessors all agree. Nothing here should be read as a claim that the largest part of the gap is understood.
Every report the tool produces prints this limitation. Do not remove it from a deliverable.
What the validation corpus actually is
engine/oracle/ holds 27 activities across 13 networks, scheduled by a real
Primavera P6 23.12 installation. It was captured from P6's own database after a
human pressed F9, by the author of cpp-cpm-engine (MIT, Dana Fitkowski). That
repository discloses the capture as fitted to its own engine; this engine has
never seen it, so for us it is genuinely held out. It is the only first-party
P6 capture held; the eleven harvested exports above are real P6 files but their
stored dates were not watched being computed.
Current agreement on the fixture: 156 of 160 field comparisons, 97.5%.
The correct reading of 97.5% is "no divergence detected on the semantics these
13 cases exercise" — never "agrees with P6". None of the 13 networks is wider
than a two- or three-activity chain. A clean sweep there is worth far less than a
clean sweep on one real 800-activity update, and the corpus that would give the
latter does not exist. engine/oracle/CORPUS-ACQUISITION.md is the plan to build
it.
All four remaining discrepancies are one representational difference that will
never agree: P6 stores a started activity's actual date in its computed start
columns, and the harness classifies them PROGRESS_ACTUAL_AS_EARLY
(engine/oracle/README.md, "The four divergences"). The free-float case on a
start-to-finish tie that this section previously listed as a fifth is now
closed: case 04 agrees 12/12, and engine/oracle/README.md Finding 3 records
that it was closed by the rework of the forward/backward bound arithmetic in
cpm.py — and that the reasoning behind that change has not been audited there.
Anyone quoting 156 of 160 should know one of those points rests on an
unaudited change. While the divergence stood it was deliberately not
patched, because the
semantics were not understood and tuning an engine to raise an agreement rate is
how the rate stops meaning anything. Agreement is a measurement; the moment it
becomes a target it stops measuring.
The harness itself carries a caveat worth reading: the XER's ERMHDR names the
exporting tool as danaf, not Primavera, so the stored dates may not be P6's.
The provenance is credible but it is not first-party.
What the corpus has already bought
One real defect, found because P6 disagreed. The backward pass and the free-float loop both took bounds from completed successors, whose late dates had been set to their actual dates — manufacturing large negative float on the predecessor of every out-of-sequence completion. Negative float is the criticality finding, so the engine was fabricating criticality from evidence that said the opposite.
Every property test held while that bug was live. The arithmetic was entirely self-consistent; it was just consistently wrong about an obligation that had already been discharged. Self-consistency cannot see that. Only an independent answer can.
That is the argument for the corpus, and it is also the reason not to trust the property tests alone.
2. Nobody has used this
No customer has run this tool on a real project. Every claim about what reviewers want comes from documents — specifications, standards, published research and one InTrans/FHWA study of eleven state DOTs — not from reviewers.
~~There is no user interface.~~ Corrected 6 September 2026. There is a
Python library, a command-line entry point, and construct serve — a local page
on 127.0.0.1 over that same command line, which computes nothing of its own.
What there is not is anything hosted: nobody can be sent a link, so trying it
means installing it first. outreach/OBJECTIONS.md §9 states it in that form.
3. ProgressMode.ACTUAL_DATES is an interpretation, not a reproduction
Primavera offers three out-of-sequence progress modes. Two of them — retained logic and progress override — have well-understood behaviour. The third does not, and the rule implemented here:
suppress a logic tie where the recorded dates already contradict it: the successor has an actual start, and the predecessor either has no actual finish or finished after that start
is this project's reading of "drives from recorded actuals rather than
forecasts", chosen because it is the reading that makes the mode distinct from
the two either side of it and produces the strict ordering
finish(override) ≤ finish(actual_dates) ≤ finish(retained).
It has not been validated against real Primavera output. The reasoning is
recorded in engine/tests/test_progress_modes.py rather than in a commit
message, so that whoever runs the P6 validation knows exactly which claim to
check.
Two consequences for a reviewer:
- A finish date computed under
ACTUAL_DATESmay not match P6's. - The upstream code this engine derives from implemented
ACTUAL_DATESas a label — it set an origin marker and suppressed nothing, producing schedules byte-identical to retained logic. That is fixed here, and a test fails if the two modes ever collapse into one again. But it is worth knowing that a mode silently behaving as a different mode is a failure this codebase has already had once.
4. Concurrency returns no governing answer under US jurisdictions
Under us_federal, us_state_or_private and neutral, the concurrency module's
governing() returns None. It reports all four doctrinal findings, states
which authority declined and why, and names the open question.
That is not a missing feature. AACE RP 29R-03 §4.2.D.1 sets out the literal and
functional theories without endorsing either, and §4.2.D treats the choice as
one for the analyst to make and defend. No document in the US precedence supplies
a rule that selects between them. A tool that picked one to avoid returning
None would be manufacturing the certainty that gets an expert excluded rather
than merely disagreed with.
Only the SCL Protocol has an encoded position, so only uk_scl produces a
governing finding — the functional / first-in-time reading, per ¶10.4 and ¶10.10.
Related limits in the same module:
- UFGS 01 32 01.00 10 and ANSI/ASCE/CI 67-17 decline as well, and each says so
in its own words. ~~Their texts are not in the corpus this module was written
from~~ — struck 6 September 2026, and it was false when written. Both texts
are held under
corpus/standards/,rules/ufgs.pyimplements more than eighty clauses of the first andrules/asce67.pyimplements all three guidelines of the second's concurrency chapter. What was true is thatforensic.concurrency's own table had no entry for either, so both fell to a message about unread documents. They were read for this purpose on 6 September: UFGS names concurrency once, at §3.8.1, as a duty to consider it, and §3.8 and §3.8.4 defer outward to ASCE 67-17; ASCE 67-17 §8.1 defines concurrency without naming the literal / functional distinction, and its §8.3 apportionment doctrine is a third position rather than a selection between the two. The verdict did not move —governing()still returnsNoneon all three US jurisdictions, correctly.engine/quality/CONCURRENCY-GOVERNANCE.mdis the record. - Float ownership is never decided. General Principle 3, AACE §4.3.E and SCL CP8 give different defaults over different quantities, and which applies is a contract question. The tool raises it and stops.
- Pacing is never decided. Of AACE §4.2.G's three tests, only the first — that a parent delay exists and precedes the pacing — is computable. The result is HYBRID or JUDGMENT and never DETERMINISTIC.
- The literal/functional-to-events/effects mapping is unsourced. AACE states the distinction as one of interval granularity; the SCL Protocol states its ¶10.3/¶10.4 distinction as one of events versus effects. The module treats them as the same axis because in practice they select the same two answers on the same facts, but they are not textually the same distinction and no source equating them has been found. A report resting on the difference should cite the paragraph, not this tool.
- Two incompatible day-count bases coexist. Concurrency day counts are calendar dates in an overlap window, inclusive of both ends. Modelled-delay impacts are working days on a named calendar. The numbers must never be added to or subtracted from one another, and nothing in the type system stops you.
5. Source Validation Protocol gaps
- Activity split and combine is not detected as itself. It is on the RP's exhaustive list of non-progress revisions, but a split produces new activity ids, so it surfaces as an addition plus a deletion. Recognising it as a split needs the mapping between old and new activities, which the schedules do not carry, and inferring one from name similarity would put a guess in the middle of a variance decomposition.
- Seven of SVP 2.1.B's eleven baseline items cannot be computed. Four are
decided from the file; the rest need the contract, the submittal log or a
person, and are emitted as
NOT_EVALUATEDorREFERRED. The checklist is honest, not complete. - De-statusing takes the progressed file's own duration as the original duration by default. The RP requires the duration that was reasonable at NTP, and nothing in the file says whether the update revised it. The tool names every activity it took the default for and refers the question — but if you ignore the referral, the reconstructed baseline is built on an unchecked assumption.
- The SVP report is only as good as the versions supplied.
python -m forensic validateruns the Source Validation Protocols over the update series you hand it; a missing revision is invisible to it. (This bullet previously said no SVP procedure was exposed through the command line.validate,concurrency,tiaandrecordare all subcommands, and METHODS.md §"Seven of the nine are reachable from the command line" listed them at the time this said otherwise.) - Bifurcation raises rather than resolves an overflow. A progress variance larger than the update period is either a revision that leaked into the half-step or a legitimate calendar-driven push into a no-work period, and the RP gives no rule to tell them apart.
6. Conformance coverage
- Fifty-five of the 103 UFGS clauses in the rule inventory are not implemented. CONFORMANCE.md §4 lists every one with its reason. Cost loading, the entire SDEF activity coding dictionary, resource loading, scheduler qualifications, submittal transmittals, SEKO minutes, narrative content and P6 admin preferences are all outside what the tool sees.
- The AACE 29R-03 and SCL packs are mostly HYBRID. Both exist — 37 and 35
rules, out of fourteen packs and 351 rules in total — but 20 of 37 and 27 of 35
need a
--termsinput, so a bareuk_sclrun reports most of its 72 rulesNOT_EVALUATED. (This bullet previously said neither pack existed and thatuk_sclreturned zero findings. Both were written, registered and tested; they were missing fromcli.py's authority map, which is now total overAuthorityand refuses at import if it is not.) - Thirty-two of the 35 ASCE 67-17 guidelines need an input the schedule does
not contain. With no
--termsfile they all reportNOT_EVALUATED. That is the standard's subject matter, not a partial implementation, but it means a bare ASCE run tells you almost nothing. - A mistyped
--termskey is ignored silently, and the resultingNOT_EVALUATEDis indistinguishable from a term you never supplied. Check keys against the tables in CONFORMANCE.md §5. - A wrong term produces a confident wrong finding. A HYBRID rule is only as
good as the number fed to it, and a
contract_completion_dateoff by a week makes every rule reading it wrong by a week, with no signal in the report. - The tool checks clauses, not compliance. A schedule can satisfy every runnable clause and still be unbuildable or wrong about the scope.
7. Methods not implemented
All nine MIPs now run, plus prospective TIA under RP 52R-06.
That is a statement about implementation, not about correctness. "Available" means the engine performs the method end to end and there are tests; it does not mean any result has been checked against Primavera, against a published expert analysis, or against a real project. Five of the nine were built in a single session and none has been run on a real schedule.
The limitation that matters is therefore not which methods exist but §1 of this document: the engine is not validated at scale. A method that runs on generated fixtures and has never met a 4,000-activity update with real progress, real calendar exceptions and real out-of-sequence work is a method whose failure modes are unknown.
Read forensic.methods.catalogue() rather than this paragraph — it reports the
same table from the code and cannot go stale the way prose does. See
METHODS.md.
8. Input formats and scheduling settings
- Three formats are read, and the format is decided by content, not by
suffix. XER, P6 XML and Microsoft Project MSPDI. The extension does not carry
the answer — both P6's XML export and an MSPDI file are
.xml— so dispatching on it gets one of the two wrong half the time;cli.pysniffs the first few hundred bytes instead. What follows from that is the real limit: a file whose opening bytes match none of the three is refused as unreadable rather than guessed at. (This bullet previously said XER only, with a.xer/.txtsuffix requirement.) - The three formats are read; they are not equally evidenced, and until
2026-09-05 nothing on a buyer-facing page said so. The corpus is 69 files
and all 69 are XER.
engine/oracle/FIDELITY.md's arm 1 puts our read of an XER against an independently written reader of the same bytes — 1,947,919 field comparisons — and there is no equivalent for the other two, because no P6 XML or MSPDI file produced by Primavera or by MS Project has ever been read by this engine. What now exists for them (arm 4) is one level weaker and its ceiling is stated with it: we write a P6 XML and an MSPDI file from a schedule we read out of an XER, MPXJ reads those files back, and the comparison establishes that an independent implementation shares our understanding of the format we write. It is not evidence that either party matches what Primavera or MS Project actually emits, and nothing in this repository can supply that. The ask, named as one: one real P6 XML export and one real MSPDI export. Until then, treat a finding drawn from a P6 XML or MSPDI submission as resting on a reader with no external witness on real output. - The P6 XML and MSPDI readers carry no resources and no costs; the XER reader
does. The two XML codecs model no resource and no assignment in either
direction — found on 2026-09-05 by an independent reader (MPXJ) of files we
write, and not by our own round trip, which compared two codecs that both omit
them and agreed perfectly about nothing.
engine/quality/CODEC-RESOURCES.mdhas the measurement and the decision. It changes no verdict: nothing in this product reads a resource off a schedule, XER included, and no conformance rule can. It is a limit on what a P6 XML or MSPDI submission carries into the engine, it is now stated in the issue log of every affected import rather than passed over (P6XML.RESOURCES_NOT_READ,MSPDI.RESOURCES_NOT_READ), and a cost or quantity question wants the XER. - The project-default lag calendar falls back to a guess. Where
LagCalendar.PROJECT_DEFAULTis in force and no project calendar id was supplied, the engine falls back to a calendar named"STD", then to the first supplied. On a multi-calendar job that silently changes every lagged link. An importer that knows the answer should say so. - The data date convention defaults to
THROUGH, which matches Primavera's treatment of an actual finish recorded on the data date. That is a compatibility argument, not a correctness one — AACE RP 10S-90 declines to settle it. The default is a choice, and it is disclosed in every report.
9. Where the code and its own documentation disagreed
Eight disagreements were found while writing this documentation. All eight have since been resolved, and the record is kept because how they arose is more useful than the fact that they did: every one was a place where a string, a comment or a default asserted something the code did not do, and none of them would have been caught by a test of behaviour.
-
collapsed_as_builtlabelled itself MIP 3.9 and performed MIP 3.8. Fixed. One network, one set of events, one collapse is modelled / subtractive / single simulation, which is 3.8; 3.9 is the period-by-period collapse. The label survived because a test asserted it —assert "3.9" in result.mip— so every run confirmed the error.tests/test_method_labels_match_the_taxonomy.pynow checks the engine's self-labels againstforensic.mip, in the forensic layer because the import contract forbidscpmcorefrom seeing the taxonomy. -
forensic.svp.REQUIRED_TERMShad drifted in both directions. Fixed.contract_scopeandcontract_termswere read and undeclared, so a reviewer had no way to know they were wanted.update_submittal_logwas declared and read by nothing, which is worse: a reviewer supplies it, nothing happens, and concludes the tool ignored their evidence. It is removed, and the check that would use it — distinguishing submitted updates from working copies, SVP 2.3.B — is recorded as unbuilt in §3 rather than implied by a table of inputs. -
The command line asserted materiality it never measured. Fixed. It marked every disclosure material, so the report printed "3 of 3 choices below would change the result on this schedule" — a sentence asserting something no computation established, in the reporting layer, which is precisely the failure this package exists to prevent.
Disclosure.materialis now tri-state;Nonemeans nobody established it, and the renderer has its own section saying so in those words.
Where it was actually fixed, because the first entry did not say and that is
how the review found it still open: Disclosure.material now defaults to
None in forensic/disclosure.py. The earlier fix changed only the CLI's
call site, so every other caller still asserted materiality by omission. An
entry that reads as closed stops the next reader checking, which makes a
half-fix more dangerous than none.
-
The conformance section's confidence label was tied to coverage. Fixed.
COMPUTED_UNVALIDATEDmeans "computed from inputs that could not be validated", which is true of every run of this command regardless of coverage, because the command does not perform the Source Validation Protocols. Coverage measures something else entirely. -
UFGS clause 073 was unaccounted for. Fixed. It is a second enforcement point for §3.12(f), which runs as
UFGS-098; running both would report one defect twice under two clause numbers, which reads to a reviewer as two problems. The pack's header now says so. -
engine/oracle/README.mdwas stale. Fixed. It now reports 156 of 160 and four named divergences, matching the harness. -
forensic.rules.asce67exported onlyREGISTRY. Fixed. ~~It now exportsRULE_COUNT,registry()andcounts_by_determinism()with the same shapes as the UFGS pack — two packs behind one interface should not need two call shapes to answer the same question.~~
Superseded 6 September 2026. The struck sentence is kept because it is
the record of what the fix was. Two of the three names it publishes no longer
exist: registry() and counts_by_determinism() were deleted from all
thirteen packs, 26 functions, having no caller anywhere under engine/src,
engine/oracle, engine/quality or demo. A pack's registry is reached as
the module constant REGISTRY, which is what forensic/cli.py builds its
authority packs from and what forensic/rules/dependence.py discovers packs
by; a pack's size is RULE_COUNT. Those two names are the pack interface,
they are asserted for every pack by engine/tests/test_pack_interface.py,
and the original point stands unchanged: thirteen packs behind one interface
do not need thirteen call shapes to answer the same question. The
per-determinism census that counts_by_determinism() returned is
Counter(r.determinism for r in PACK.REGISTRY.rules) at any caller that
wants it. Recorded in engine/quality/UNREACHABLE.md, section 7.
forensic.mip.candidates()andall_profiles()were public and missing from__all__. Fixed.
10. What none of this is
Nothing produced by this tool is legal advice, and nothing in it is advice about any project. The concurrency and jurisdiction modules encode what published documents say, with the paragraph references so a reader can check, and leave the places they disagree disagreeing.
The FRE 902(13)/(14) certification the evidence module produces is draft text for a qualified person to review, adopt and sign. It is not a certification until a person with knowledge signs it, and this engine cannot be that person. Whether any given signatory is "qualified" within the Rule is a judgement for counsel. Self-authentication addresses authenticity only, not hearsay — a self-authenticated record still has to get past FRE 802 — and Rule 902(11)'s procedure, which (13) and (14) incorporate, requires reasonable written notice to the opponent and an opportunity to inspect. Missing that step forfeits the benefit.
Gap detection states observable facts about the set of documents provided and nothing more. It does not assert that anything was lost, destroyed, withheld or concealed, because FRCP 37(e) makes that a question of preservation duty, reasonable steps and — for the severe sanctions of 37(e)(2) — intent to deprive, none of which is visible in a set of files. The most common innocent explanation for a missing update is that nobody sent it.
11. What the delay-injection matrix demonstrated a method cannot do
Added 2026-09-02. Every other oracle tier in this repository measures the
scheduler against P6 or the reader against MPXJ. engine/oracle/injection/
measures the forensic layer directly: it injects a delay of hand-counted size
into a five-activity network and asks whether each MIP reports the delay that was
injected. engine/oracle/injection/MATRIX.md carries the run, the predictions
registered before it, and the mutations that verify it can fail.
All five exercised methods recovered every headline figure, including a placebo (nothing injected, every method reported zero) and a neutral control (a float activity extended five working days against fifteen days of float, every method reported zero). The limitations below are what the same run showed about attribution, and each has a demonstration behind it rather than a reservation.
MIP 3.3 stops explaining a delay at the point the delay becomes decisive. Two
seven-day slips on the same fixture: where a critical activity was extended, the
windows analysis attributed every day to a named cause on a named activity; where
a float activity overran far enough to take over the driving path, it
attributed nothing — the whole movement landed in the residual PATH_SWITCH /
UNEXPLAINED. The magnitude is right in both. So explained_fraction reading
0.0 can mean the analysis did no work and can mean the delay was large enough
to change the critical path, and this method does not distinguish them. A report
quoting an unattributed figure has to say which the reader is looking at, and no
method here can tell them.
MIP 3.1's per-activity table counts float consumption as a delayed finish. On
the neutral control the project figure is correctly zero and
candidate_as_built_critical_path is correctly empty — exceeds_planned_float
is doing the RP's late-date test, as engine/quality/MIP-COVERAGE.md §2 records.
But delayed_finish and extended_duration are both True for the float
activity, so a reader summing delayed_activities obtains five days of delay
on a project that was not delayed. The headline is right; the table beneath it
is a list of movements, not of delays, and is not labelled as one.
MIP 3.2's per-period variance was not the delay that accrued in that period.
Fixed 2026-09-02; the reservation is withdrawn. On a seven-day delay the two
periods published +21 and −14. They summed to the gross figure and RP §3.2's
reconciliation held; what a reader was handed was twenty-one days of variance in
a period ending before any of it happened. Activities were filed into a period by
their as-planned finish while the level was taken over their as-built
finishes, so any float activity overrunning its period contributed the whole
overrun to the earlier period and a negative correction to the later one — and
Period.variance_days reads a negative as time recovered, which on that fixture
nothing was.
An activity is now filed on the later of its two finishes, so a period's
figure can only draw on dates its own boundary has passed. The same injection
publishes [0, +7]. Filing on the as-built finish alone was rejected: it has the
same defect with the signs exchanged — an activity planned late and finished
early would file into an early period and drag its planned date forward out of
it, claiming a recovery weeks before the work it was recovered against was due.
A genuine recovery remains representable and is tested for
(engine/tests/test_period_filing.py), which reads the per-period figures:
MATRIX.md §6 records that a check reading only the total cannot see a filing
defect. MATRIX.md §4.2 still describes this as an unfixed defect candidate;
that section is the record of the finding and is superseded here.
A constraint-driven delay has no MIP 3.6 input form at all.
cpmcore.modelled.DelayEvent attaches to the network as a Finish-Start
predecessor and carries no constraint field, so a delay imposed by a date
constraint — ten working days on the fixture, recovered in full by all four
observational methods — cannot be handed to the impacted-as-planned method in any
form. That is n/a rather than a miss, and it bounds what a modelled analysis
can be asked about on a schedule whose delays live in its constraints.
What the matrix does not establish is in MATRIX.md §8 and is short enough to
repeat: five activities, one injection at a time, no concurrency, no fuzz, and no
attribution of cause to a party anywhere — the engine publishes
attributes_cause: false and means it. MIP 3.5 was not exercised; §7 there
says why.
Superseded on 2 September 2026. That paragraph also said 3.7, 3.8 and 3.9
were not exercised, and that the matrix had no calendar mix. Both were true when
written and are no longer: the collapse methods were injected into the first
fixture in this repository that satisfies §3.8.E.1, on a network carrying a
five-day and a seven-day calendar, and engine/oracle/injection/MATRIX-COLLAPSE.md
has the runs. The four bounds below come from it.
11a. The collapsed as-built — four bounds, each demonstrated
The total depends on an anchor nobody states. Both sides of a collapse are
floored at a data date. MIP 3.8 infers one from the earliest actual start; MIP
3.9 cumulative uses the first base's. On the reference fixture the same model
and the same three delay events give fourteen working days at one anchor and
eleven at the other, and the concurrency figure moves with it — plus five
against plus three on a second run. Ask which anchor a collapse figure was
produced at. SubtractiveMultiBaseResult.anchor_note states it in the record.
A small delay near a calendar boundary can be recovered as zero, correctly. One working day of delay on a five-day-calendar activity feeding a seven-day one is recovered as one day at the inferred anchor and zero at the hard anchor — because the project finish genuinely did not move; the day landed inside a weekend the successor was waiting through anyway. Do not read a zero as the method having missed something. It also means a dose curve for MIP 3.8 is a property of the anchor as much as of the method.
A delay recorded in line, with no §3.8.K.2.c redundant tie around it, produces a collapse larger than the delay removed. Thirty-three and thirty-four working days out of eighteen days extracted, with a successor finishing before its own predecessor — and §3.8.E.1 still reads PASS. The only disclosure is the severed-chain table, which names every broken pair and is correct. Read that table before the total: if it is not empty, the total is not a delay figure.
A per-event figure is not an apportionment. Lengthening one delay changes another's reported contribution without that second delay being touched — on the fixture, one event reads four days and then three. A single-event collapse is a statement about that event in that programme, not a property of the event, and the per-event figures do not sum to anything a reader should quote.
The bound on all four. These come from fixtures built to satisfy §3.8.E.1, with their dates copied out of the same scheduler then asked to reproduce them. A PASS is evidence about the method, never about the population. Of the sixty-nine corpus files, sixty-six schedule; of those, forty-eight record no actual dates at all and eighteen record them and fail. None passes.
One thing the collapse family cannot have, and it is a fact about testability rather than a defect. There is no clean placebo on a mixed-calendar as-built. Dosing every delay to zero and extracting them still moves the finish by one day, because a recorded zero-duration activity is still a calendar boundary — its seven-day successor waits for the Monday its five-day calendar imposes. The finish really did move, so the method is right. It means the cheapest control in the whole design is unavailable here, and the placebo that is available has to be read with that in mind.
11b. The method is an input to the answer, and here is what that costs
Every honest account of forensic schedule analysis says the choice of method affects the result. It is usually said as a caution. It can now be said as a number, on one fixture, on one day, with both figures hand-counted.
The same delay event is worth nothing and worth a week, depending only on
which method is asked. DLY-APPROVAL — an owner-caused approval delay that
was real, was recorded, and never drove the project finish — is worth:
- 0 working days under the subtractive collapse of MIP 3.8. Remove it from the as-built and the finish does not move, because a competing path was carrying the critical route the whole time.
- 5 working days under the additive impact of MIP 3.7. Insert it into the contemporaneous base that did not yet contain it, and the finish moves five days, because at that update the competing path had not yet absorbed it.
Neither figure is wrong and neither method is misapplied. They answer different questions: what would have happened if this had not occurred against what happened when this was introduced. On a schedule where the driving path changes hands — which is most real schedules — those questions have different answers, and the difference is not small. Here it is the whole of the claim.
What follows for anyone quoting a delay figure from this tool. A number
without its method is not an answer. The engine publishes the method with every
figure and attributes_cause: false alongside it, and neither of those is
boilerplate: this is the case that shows what they are protecting against. If an
expert on the other side runs a different MIP and gets a different number, that
is not necessarily either party being wrong, and a report that does not say so
is inviting a cross-examination it cannot survive.
Two more from the same run, recorded because they bound what MIP 3.7 can be
asked. Its dose curve terminates — at 22 event-days on this fixture, where
one event's duration grows past the next event's onset, §3.7.E.7 fires and the
method withholds the total rather than publishing one. That is a correct
refusal. And the double-count risk of impacting an event into a base that
already contains it does not reach any figure 3.7 publishes: both sides of
the subtraction move together, so combined_days and the total hold at 6 and 14.
Only the finish moves, 2026-05-12 to 2026-05-20 — and §3.7.E.14 already
declares that finish is a figure the method cannot check.
12. Output surfaces that are still hard to read aloud
Audited whole on 2 September 2026, for the first time
(docs/ACCESSIBILITY.md, which owns every finding and is the page to read; this
entry exists so that "everything the tool cannot do" does not silently omit the
one axis its own author reads on). The Markdown report, the HTML report and the
JSON record came out sound. What did not, and is recorded there rather than
here because it moves: two of the eight error and refusal paths do not read as
sentences, and one warning in forensic.evidence verify dumps a Python repr
at the reader. That last one is a stated tension and not an oversight —
accessibility and stating evidence precisely pull against each other there, and
it was left visible rather than resolved by quietly dropping the precise form.
This matters commercially and not only ethically: a forensic report is read aloud, quoted in a letter and pasted into a filing, so text that works only visually is broken output rather than ugly output.
13. Every check that needs a second machine has never run
Recorded overnight 5-6 September 2026. This entry is here because the rest of the document describes checks that execute, and the checks in this section do not. A check that has never run is not weaker evidence than a check that has. It is not evidence.
Continuous integration has been dead since 3 September and 284 commits have
landed since. .github/workflows/ci.yml defines five jobs — check,
differential, package, record and reproduce — and it is the half of the
verification that needs real interpreters, real runners or a network, which is
to say the half nobody would type by hand. GitHub reports 429 workflow runs
for this repository. Of the 200 most recent, none concluded success; the
oldest of those 200 was created at 2026-09-03T09:04:05Z, so the failing streak
covers every run in more than two days. Runs on today's commits conclude
failure in four to seven seconds, which is the shape of a job that never
started a step. engine/quality/CI-AUDIT.md §0a has the cause in GitHub's own
words: "the job was not started because recent account payments have failed or
your spending limit needs to be increased." It is a billing state and it is the
one defect in this repository that cannot be fixed from inside the tree.
The reproduction matrix has never executed once, in the life of the
repository. reproduce is declared needs: record, and record has never
passed, so the matrix has been skipped on every run of its existence
(CI-AUDIT.md §1, which measured that before the billing halt and again after
it). That matrix is five legs — Ubuntu and Windows against Python 3.11, 3.12 and
3.13, less the leg that writes the record — and it is the only place two legs of
this project's central claim are measured at all:
- Cross-version determinism is unknown, not established.
engine/quality/REPRODUCIBILITY.md§2 says so in those words.pyproject.tomladvertisesrequires-python = ">=3.11"; nothing has ever compared an answer computed on 3.11 with the same answer computed on 3.12 or 3.13. - Cross-platform determinism is assumed. What is measured is narrower than
it sounds and is worth stating exactly, because the broader sentence is the
one that would travel: the answer digest is identical across five fresh
processes under five explicit
PYTHONHASHSEEDvalues, one of which disables randomisation and is the control — on one machine, one operating system and one interpreter, Windows 11 running CPython 3.13 (engine/quality/DETERMINISM.md§6, which records thatpy --listreports a single interpreter on that machine). Five seeds on one Windows box is not two platforms. One real platform difference has been found rather than assumed away —core.autocrlfgives the committed.xerdifferent bytes on Windows and the same parse — and it is measured by a unit test rather than by the matrix (REPRODUCIBILITY.md§3).
The three-way differential against MPXJ has never run automatically either
(CI-AUDIT.md §1). It is the only check in this project that can say we are
wrong without asking Primavera, and it has only ever run when somebody typed it.
There is no install route but a clone. pip install construct-engine and
pipx install construct-engine resolve to nothing; the name is unregistered,
and docs/RELEASING.md §9 records the date and the HTTP status of that check.
The package job — which builds the wheel and installs it the way a stranger
would, with --no-index — passed exactly once, on 3 September, before the halt.
docs/INSTALL.md states the clone requirement on its first screen and is the
page to read; it is repeated here because "what this cannot do" should not need
a reader to have already found the installation page.
What this does not license anyone to say. Nothing above means the engine is
non-deterministic, or broken on 3.12, or unusable on Linux. It means nobody
knows, and the difference between unmeasured and fine is the whole subject
of this document. The gate that a developer can run — engine/check.py, the
unit suite, the oracle suite, the import contracts and the P6 agreement — is
unaffected by all of this and runs on every machine that has a clone.
What this tool may not be used to produce
A limit that is not ours and cannot be engineered away. UFGS 01 32 17.00 20 (NAVFAC, 05/25) makes third-party processing of an XER cause for rejection. The submission a contractor sends must come from their own Primavera, not from anything this tool wrote.
That forecloses a whole class of feature permanently, on exactly the federal work
this product targets: no corrected file, no re-export, no "fixed" schedule handed
back, however carefully any of it were built. cpmcore ships three writers
(write_xer, write_p6xml, write_mspdi) and no shipped command reaches any
of them -- an accident of history that turns out to match the constraint, and
which should stay that way unless someone can name a use outside the differential
oracle that this clause permits.
It also bounds the honest shape of anything contractor-facing. respond.py
groups a report's findings by what the recipient must do and names the missing
inputs an abstention depends on; it produces prose, never a file, and
docs/CONTRACTOR-MODE.md records why that is the ceiling rather than the first
version.
Recorded here 2026-09-02. The clause was established in
corpus/analysis/FIELD-EVIDENCE.md and cited in business/ and
docs/CONTRACTOR-MODE.md, and was absent from the page that owns limitations --
found by an agent that grepped this file for it while designing against it.
Source: docs/LIMITATIONS.md. Source commit date: 2026-09-06.