AI Skill Report Card

Documenting Mathematical Case Records

A-85·Oct 6, 2026·Source: Extension-page
14 / 15

A case record answers one question precisely: "What do we know about case n, and exactly how well do we know it?"

Minimal skeleton for one case:

Markdown
# Case n = 11
Recommendation▾
Consider trimming the 'Best Practices' and 'Common Pitfalls' sections since they substantially overlap with content already stated in the numbered Workflow steps—could be merged to tighten the skill.
  • Best known packing: side = 3.877083590022814... (Trump, 1979) [Proven]
  • Lower bound: 3.877083590022814... (Queuingtheorydotcom 2026, confirmed T-060)
  • Upper bound: 3.877083590022814... (Trump 1979, confirmed T-011)
  • Gap: 0 — Solved

Upper bound (verified)

...exact value, source, evidence ID...

Lower bound (verified)

...exact value, source, evidence ID...

  • T-xxx S# V# C# — author · date · claim, composition, next rung, significance, novelty

...open questions, caveats, falsified arguments, what remains to be shown...


Every numeric claim gets: **exact form (closed form or minimal polynomial) → who proved it → verification tier → evidence link**. Never report a rounded decimal as if it were the definition of the value.
15 / 15
Progress:
- [ ] Gather every result (register entries) that bears on this case
- [ ] Determine current best lower bound and best upper bound, each with exact form
- [ ] Assign/confirm verification tier (S/V/C) for each contributing result
- [ ] Compute the gap; mark solved only if bounds meet exactly
- [ ] Write the visual/numeric summary block (packing + number line description)
- [ ] Write each bound entry: value, prover, kind, scope, caveats, evidence ID
- [ ] Write each register result entry: claim, composition, next rung, significance, novelty
- [ ] Write the case file's own narrative account (history, corrections, context)
- [ ] Cross-check citations and evidence IDs resolve
- [ ] Render to the case's own address/page and confirm it also surfaces in index/survey/atlas

1. Establish the bound ladder

For each case, identify:

  • Reported values (as published, possibly only numerically verified)
  • Verified values (independently re-derived/checked, with exact algebraic witness)
  • The gap = upper − lower. Zero gap (exact value) ⇒ solved, render in accent/highlight.

2. Tier every claim with a verification ladder

Use an explicit rung system, e.g. S# V# C#:

  • S (Significance tier): how central/novel the result is
  • V (Verification depth): 0 = unverified claim, increasing with independent replay, adversarial review, formal proof
  • C (Confirmation count): how many independent confirmations exist at that depth

State explicitly what the next rung requires ("needs two adversarial AI reviews by distinct reviewers and a human oversight record"). Never silently round up a tier — if cached audits are stale or a public verification route doesn't currently re-confirm a result, say so and state what tier it actually stands at, including any history of regression ("held V4/C5 until the ladder change of 2026-09-30").

3. Report exact forms, not approximations

  • Prefer closed form (e.g. 2 + 4/√5) or minimal polynomial with explicit degree and coefficients.
  • Rounded display digits are a convenience for the reader, never the definition of the endpoint — say so explicitly if there's any risk of conflation.
  • When a bound is a root of a polynomial, state the polynomial, its degree, and (if needed) the isolating interval or inequality picking out the right root.

4. Record the composition of derived bounds

When a bound is derived (not primary), show the derivation chain explicitly:

  • What is scaled/invariant/preserved under the transformation
  • What lemma supplies the new containment/exclusion
  • What limit or density argument (if any) extends a strict inequality to the boundary
  • Which steps are mechanical corollaries vs. which required new proof

5. Separate "what was found" from "what the method cannot do"

Every strong result should state its own ceiling honestly:

  • What the method structurally cannot reach (e.g., "a certificate cannot exist above ceil(√n)·B")
  • Where two independent routes (e.g., search vs. rejection) each fail to close, and why — distinguish a stalled search from a structural limit of the method.
  • Numeric convergence across independent runs is evidence, not proof: say so.

6. Write the narrative honestly, including falsifications

If a historical published argument was checked and found not to close, say exactly that — cite the gap, state that the printed argument is not relied upon, and show the independent repair/replacement that restores the result. Do not quietly inherit a citation's authority without re-verification.

7. Novelty labeling

Classify each result's novelty explicitly and narrowly:

  • previously-published — present in an identified source
  • apparently-novel — not found in the recorded search, subject to stated search-scope gaps Always state the scope of the search that grounds a novelty claim; "apparently novel" is a claim about search coverage, not absolute priority.

8. Cross-render consistently

A case record is written once as a source file (e.g. packing/frontier/n-NNN.md) and rendered to multiple surfaces (standalone page, survey list, atlas/grid overview, index). Keep all case-identifying facts (bound values, tiers, citations) in the single source record; downstream pages must not duplicate or fork the data.

Recommendation▾
Add a brief example of a 'bad' output (e.g., a poorly tiered claim or rounded-decimal mistake) alongside the good examples to reinforce contrast per best-practice guidance.
18 / 20

Example 1 — Solved case: Input: Case n=11, upper bound from Trump 1979 (3.877083590022814...), new lower bound from Queuingtheorydotcom 2026 proving global optimality via counting argument. Output:

Best known packing: 3.877083590022814... (Walter Trump, 1979, Proven)
Verified upper bound: 3.877083590022814... [T-011, V3/C3]
Verified lower bound: 3.877083590022814... [T-060-derived, V3/C3, Evidence E-n011-global-optimality-independent]
Gap: 0 — Solved
Note: cached audit digests are stale; public RUN_ALL does not re-verify them;
this repository independently replayed inputs and composed exact obligations.

Example 2 — Open case with a falsified historical argument repaired: Input: Stromquist 1984/2003 Figure 14 point set is cited as proving s(11) ≥ 2 + 4/√5, but the printed argument doesn't actually close. Output:

Claim: s(11) ≥ 2 + 4/√5, by a source-distinct repair of Stromquist's Figure 14 set
(replacement point G' = (79/100, 37/20) restores localization and capacity chain).
Note: The printed 1984/2003 argument does not close; it is not relied on.
Next rung: needs two adversarial AI reviews + human oversight; a second,
method-distinct verification (pose-space interval audit) would sit beside this rung.
Significance: resolves proof status of a bound cited as proved for ~40 years.

Example 3 — Reporting a structural ceiling, not just a stalled search: Input: A covering-LP search stalls at container side 3.82 for n=11; need to distinguish method limit from search failure. Output:

A certificate needs mass strictly below n; neither tested site set reaches it.
Ceiling check: a certificate cannot exist above ceil(√n)·B = 3.9908 for n=11,
so 3.82 is not the method's ceiling — there is 0.1808 of runway above it.
If τ*(3.82) = exactly n, this is the one point where neither the certificate
route nor the rejection route can close — a structural limit, not a stalled search.
Recommendation▾
The domain is quite niche (packing problems specifically); consider a one-line note on how to generalize the S/V/C tier system to other mathematical domains for broader applicability.
  • Always pair a numeric value with its exact algebraic form (fraction, surd, or minimal polynomial) the first time it's introduced.
  • State verification tier for every claim, including historical/inherited ones — do not let a tier be assumed from a result's age or fame.
  • Explicitly name what is required to advance to the next tier.
  • When two independent methods each stop short of closing a case, explain why each stopped and whether that's a search limitation or a structural one.
  • Distinguish "reported" from "verified" throughout; a value can be reported long before anyone verifies it.
  • Keep narrative honest about priority — "apparently novel" always comes with the scope of the search behind it.
  • Preserve retained-but-superseded bounds in the record (e.g., calibration rungs, earlier weaker bounds) rather than deleting history.
  • Do not present a rounded decimal as the exact value — always carry the closed form or polynomial alongside it.
  • Do not inherit a tier (e.g., "proved since 1979") without checking whether the printed argument actually closes; if it's broken, say so and cite the repair.
  • Do not conflate numeric convergence of independent search runs with a proof — label it as measurement/evidence.
  • Do not claim a search method has "run out of room" without checking the method's actual structural ceiling.
  • Do not silently fork case data between the index, survey, and standalone page — one canonical source record per case.
  • Do not state "apparently novel" without bounding the claim by the scope of the search that supports it.
0
Grade A-AI Skill Framework
Scorecard
Criteria Breakdown
Quick Start
14/15
Workflow
15/15
Examples
18/20
Completeness
19/20
Format
14/15
Conciseness
13/15