Analyzing Diagram Delivery Reliability
YAML--- name: analyzing-diagram-delivery-reliability description: Analyzes AI-generated diagram delivery pipelines using request-level telemetry to distinguish "promised vs. accepted artifact" outcomes, and writes editorial research reports on reliability findings. Use when investigating failure rates in AI content-generation systems, auditing delivery logs against a shared denominator, or writing technical research posts that report metrics without overclaiming correctness. ---
Given raw request logs for an AI generation feature (e.g., diagrams, code snippets, images), compute delivery reliability using a shared denominator—not per-stage percentages that hide drop-off.
Denominator: unique request IDs where the system promised an artifact
Numerator: requests that reached an accepted artifact (rendered, validated, delivered)
Failure rate = 1 - (accepted / promised)
Example:
- 4,870 unique diagram-promised requests over 7 days
- 3,294 reached an accepted artifact
- 1,576 did not (32.4% failure rate)
Report this as: "In 4,870 diagram-promised requests, 3,294 reached an accepted artifact and 1,576 did not." Never collapse this into a single misleading percentage without the raw counts alongside it.
Progress:
- Define the denominator: what event marks "promised"?
- Define the numerator: what event marks "accepted"?
- Pull unique request IDs over a fixed observation window (e.g., 7 full days)
- Compute counts and rate; verify counts sum correctly (numerator + gap = denominator)
- Identify what the metric does NOT certify (e.g., delivery ≠ correctness)
- Draft headline finding using exact counts, not rounded percentages alone
- Write byline-style report: finding → method → scope limitation
- Render any supporting chart deterministically (SVG → PNG) from the aggregate data, never per-user detail
Step 1: Define promised vs. accepted precisely. "Promised" = the system committed to producing an artifact (e.g., user requested a diagram and got an acknowledgment/render attempt). "Accepted" = an artifact existed that passed minimum acceptance criteria (rendered without error, met schema, was returned to user). Document the exact log event names used for each.
Step 2: Use unique request IDs as the denominator. Never use raw event counts (which double-count retries) or session counts (which hide multi-request sessions). One request ID = one measurement.
Step 3: Report absolute counts alongside the rate. A "32.4% failure rate" alone is easy to misread. Always pair it with "1,576 of 4,870" so readers can verify the math and sense the scale.
Step 4: State what the metric does and does not prove. Delivery (an artifact was returned) is not the same as correctness (the artifact was engineering-accurate). Explicitly separate these in the writeup — this is the core editorial discipline that distinguishes rigorous reporting from marketing claims.
Step 5: Use privacy-safe aggregates only. Work from counts and rates, never individual user content or request payloads, when the data will be published.
Example 1: Input: Logs show 4,870 unique diagram-request IDs in a 7-day window; 3,294 have a corresponding "artifact_accepted" event. Output: "AI single-line diagram reliability: what delivery logs now prove — In 4,870 diagram-promised requests, 3,294 reached an accepted artifact and 1,576 did not (32.4% unresolved). The shared request denominator measures delivery—but still cannot certify engineering correctness."
Example 2: Input: A stakeholder asks "what's our diagram success rate?" citing only "68%." Output: Reframe as: "68% (3,294/4,870) of diagram-promised requests reached an accepted artifact over 7 days. This measures delivery completion, not whether the diagram was technically correct — that requires a separate accuracy audit."
- Always name the observation window explicitly (e.g., "seven full days") — partial days skew rates.
- Pick a denominator once and hold it constant across all reported metrics in the same piece.
- Pair every percentage with its raw numerator/denominator.
- Separate "delivery reliability" claims from "correctness/quality" claims — never let one imply the other.
- Use deterministic rendering (fixed SVG → PNG pipeline) for any published chart so figures are reproducible from the same aggregate data.
- Attribute the piece (byline, date) as editorial research, signaling it's an analysis, not a vendor benchmark.
- Don't report percentage-only stats without underlying counts.
- Don't switch denominators mid-analysis (e.g., unique requests vs. total events) — this silently changes what the number means.
- Don't imply "delivered" means "correct" — always flag that engineering correctness needs separate verification.
- Don't use session-level or user-level counts as a substitute for request-level IDs — this over- or under-counts activity.
- Don't publish per-user or per-request raw content in aggregate reporting — stick to counts and rates.