DILIGENCE / LIMITS

What PlotWarden does not claim.

The site is explicit about boundaries so an intended architecture is not mistaken for achieved reliability.

It does not eliminate human judgment.

Ambiguous conflicts, novel risk, and value decisions still require people. The design minimizes routine intervention; it does not automate genuine adjudication away.

It does not make untrusted content safe by naming it evidence.

Provenance and policy reduce risk, but source content can be wrong, adversarial, or stale. The system must preserve the distinction between data and instruction throughout ingestion and execution.

It does not prove the 10,000-captain goal.

That statement is ratified intent for Black Skies. Only dated benchmark evidence can establish achieved performance.

It does not claim a permanent competitive gap.

The agentic development market changes quickly. Comparisons are dated, definition-bound observations with caveats. The project should change when the market invalidates part of its thesis.

A public limit is part of the evidence model, not a disclaimer appended after the pitch.

Five unresolved system risks remain visible.

Graph completeness, authority-policy quality, poisoned-evidence containment, reconciliation correctness, and operating cost remain open until dated, reproducible artifacts can bound them.

  • Graph completeness — unknown dependencies can sit outside a computed frontier; 'dependency unknown' is therefore a designed control-plane result that taints conservatively instead of inventing precision.
  • Policy quality — a permitted action can still reflect a bad policy.
  • Poisoned evidence — the poison epoch fences known consumers while unknown consumers remain undiscovered until appended.
  • Reconciliation correctness — preserving work safely requires false-preserve, false-recall, determinism, and concurrency tests.
  • Operating cost — local and remote execution economics require measured workloads, not assumptions.

Open evidence questions.

These remain questions until a dated artifact supplies a reproducible answer.

  1. How accurately does trace-first acquisition recover dependencies in a substantial repository?
  2. Which changes produce false preserves or false recalls at the impact frontier?
  3. Can unaffected work be preserved safely under concurrent ratified changes?
  4. Are the public work dispositions deterministic and reproducible?
  5. How do conflicting changes, repeated repairs, budget exhaustion, and non-converging plans stop?
  6. How reliably are poisoned claims and every known dependent consumer discovered?
  7. What evidence proves sealed orders and commitment gates prevent scope or authority bypass?
  8. What direct evidence quantifies manual reconciliation frequency and cost in enterprise agentic development?
  9. What human-adjudication rate is compatible with minimal human intervention by design?
  10. What ledger storage, retention, redaction, and recovery policies are required?

Evidence posture.

Intended behavior, demonstrated records, and unproven claims remain separate.

Unproven

Reliability, operating cost, graph completeness, and differentiation remain open evidence questions rather than achieved outcomes.

Inspect the record