Skip to content

project-estimation: add calibrated delivery estimates - #134

Merged
sumitake merged 17 commits into
mainfrom
dev/codex/project-estimation
Aug 22, 2026
Merged

project-estimation: add calibrated delivery estimates#134
sumitake merged 17 commits into
mainfrom
dev/codex/project-estimation

Conversation

@sumitake

@sumitake sumitake commented Aug 21, 2026

Copy link
Copy Markdown
Owner

Summary

  • Adds the generated project-estimation skill and deterministic offline estimator/reconciler.
  • Auto-invokes the estimate checkpoint from formal architecture, orchestration, and teamwork planning paths.
  • Reports focused agent wall-clock and API-equivalent token cost as headlines, with marginal cash, pricing freshness, and quota delay available separately.
  • Adds explicit staged bootstrap and promoted calibration contracts, closed public schemas, maintenance verification, archive membership, README, and architecture documentation for unreleased source version 6.2.0.

Linked workspace source

  • Workspace PR: https://github.com/sumitake/agent-collab-workspace/pull/2921
  • Workspace source merged first as signed GitHub squash commit ebf821c60cb60c4aff566c55f08e25382cd412d7.
  • The committed bootstrap bytes match the merged workspace evidence manifest: aggregate 1d19079c…, pricing fdc67cab…, quota cc0aabf0…, and notification 0b098d20…; receipt 9d104bacdbc251fd07a441b071b3111c6d274c07e1feddc18797ebf13eb22449.

Validation

  • Exact candidate head: af5e68fdb03fffb8d286ab089184ca8d88281357 over base c08e7efde157b67777cc231df978b8f62fd5f87b.
  • All 17 feature commits have valid signatures from the repository owner; the replacement head was independently verified with git verify-commit.
  • Public tests: 264/264 passed.
  • Scripts discovery: 448/448 passed.
  • Generated skills and marketplace are byte-current.
  • Maintenance admission, release consistency, active-tree public-export safety, Python 3.10 API scan, canonical archive build/readback, and git diff --check pass.
  • Fresh remote CI on the exact head is green across every required check.
  • Direct independent xAI/Grok exact-head review: APPROVE, confidence HIGH, with no actionable findings. Durable receipt: project-estimation: add calibrated delivery estimates #134 (comment) (ses_fd813746dffeBci1uIwDa7PESw).
  • Exhaustive current-head review inventory: project-estimation: add calibrated delivery estimates #134 (comment); 0 actionable findings and no unresolved review thread at cutoff 2026-08-22T05:30:35Z.
  • The production bootstrap is enhancement-duration-only and descriptive. Greenfield returns typed unavailable; token/API-equivalent-cost, marginal-cash, quota-delay, wait, and rework evidence remain explicitly unavailable rather than being imputed as zero.

Required checklist

  • Trigger verification
  • Version bump to 6.2.0 source surfaces
  • Changelog fragment
  • Local validation passed

Governance and release boundary

This remains a source-only PR. The exact reviewed head, fresh CI, and complete current-head review-surface inventory are satisfied. No tag, release, install, activation, restart, or canary is authorized.

author: Codex (OpenAI family), acting for repository owner John Osumi
standing_directives: public-source boundaries, signed commits, validation, independent review, privacy, and fail-closed release gating followed
tier: 3
cross_check: PROCEED — direct non-coordinator xAI/Grok review of exact base c08e7ef and head af5e68f returned APPROVE with HIGH confidence and no actionable findings; session ses_fd813746dffeBci1uIwDa7PESw; durable receipt #134 (comment); the operator expressly authorized this reviewer route and coordinator bypass, and no consumed operation was replayed or relabeled
post_condition: exact-head independent review, local validation, fresh remote CI, and the exhaustive current-head review inventory are complete; normal operator-authorized source merge is the only remaining gate; tag, release, install, restart, activation, and canary are prohibited
mcp_coverage_gap: NONE
contributor_rights: OWNER-AUTHORED
operator_reserved: yes - the operator authorized governed source work and merge only after this PR's own gates; every release and activation action remains separately controlled and currently prohibited

@coderabbitai

coderabbitai Bot commented Aug 21, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: fef59646-7b99-43bd-bfa8-a6b8c66eaf4d


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment thread scripts/test_verify_project_estimation_maintenance.py Fixed
Comment thread scripts/test_verify_project_estimation_maintenance.py Fixed
Comment thread scripts/test_verify_project_estimation_maintenance.py Fixed
@sumitake
sumitake force-pushed the dev/codex/project-estimation branch from ed34737 to b11e3c5 Compare August 22, 2026 04:43
@sumitake

Copy link
Copy Markdown
Owner Author

Independent exact-head review receipt

  • Route: direct non-coordinator reviewer, as expressly authorized by the operator after the coordinator path proved unavailable
  • Reviewer family/model: xAI / opencode-go/grok-4.5, effort high
  • Session: ses_fd813746dffeBci1uIwDa7PESw
  • Base: c08e7efde157b67777cc231df978b8f62fd5f87b
  • Head: af5e68fdb03fffb8d286ab089184ca8d88281357
  • Repository posture: detached sealed clone at the exact head, remote removed, clean before and after review
  • Terminal verdict: APPROVE
  • Confidence: HIGH
  • Findings: 0 Critical, 0 Important, 0 Minor; no actionable findings

The reviewer independently confirmed the three prior-head remediations: material-unpriced evidence remains semantically admissible but fail-closed at publication with distinct diagnostics and a direct regression; omitted artifact_scope_hash is derived while a supplied mismatch is rejected; and micro-USD cost quantiles use the dedicated nullable_microusd schema definition.

Reviewer-executed verification included 113 focused estimator/maintenance tests, 68 project-estimation discovery tests, 15 archive/security-contract tests, release consistency, exact scope-hash probes, material-unpriced release-block probes, and schema-reference probes. All completed successfully.

Observed direct-review usage: 115,744 input tokens, 4,288 output tokens, 3,445 reasoning tokens, 595,840 cache-read tokens; API-equivalent reviewer cost recorded by the runtime as $0.575806.

This receipt is source-governance evidence only. It does not authorize a tag, release, publication, installation, activation, restart, or canary.

@sumitake
sumitake marked this pull request as ready for review August 22, 2026 05:28
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@sumitake

Copy link
Copy Markdown
Owner Author

Final current-head review inventory

Every available review surface was paginated and adjudicated:

  • Formal reviews: one historical CodeQL COMMENTED review on an earlier commit. Its three chmod findings are all resolved and outdated; fresh exact-head CodeQL is successful with “No new alerts in code changed by this pull request.”
  • Inline review comments/threads: three total; all resolved and outdated; none remain actionable.
  • Issue conversation: CodeRabbit reports that this OSS repository requires a manual review; the direct xAI/Grok exact-head approval supplies that independent manual review. The Codex security-review connector separately reported a usage-limit condition; it is neutral/unavailable optional evidence, contains no finding, and was neither retried nor used as approval. The direct Grok receipt is recorded above.
  • Checks: every required check is successful. The sole skipped job is the non-applicable Dependabot ARM job.
  • Check annotations: the 32 failure-level annotations on successful Python/frontmatter jobs are deliberate output from fail-closed Dependabot negative tests. The assertions passed in all jobs, so they are informational test evidence rather than findings on this PR.

Disposition at cutoff: 0 actionable findings; 3 historical findings resolved/outdated; no unresolved review thread; exact-head independent review approved with high confidence. The PR is eligible for the already-authorized normal source merge after final head/check/signature readback. This does not authorize any tag, release, publication, installation, activation, restart, or canary.

@sumitake
sumitake merged commit 398b811 into main Aug 22, 2026
21 checks passed
@sumitake
sumitake deleted the dev/codex/project-estimation branch August 22, 2026 05:32

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: af5e68fdb0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


def reconcile(prior_result: Mapping[str, object], actual: Mapping[str, object], pricing: Mapping[str, object]) -> dict[str, object]:
prior = validate_result(prior_result); actual2 = validate_actual(actual); pricing2 = validate_pricing(pricing)
if prior["result_kind"] != "estimate" or "detail" not in prior: raise EstimationError("prior result must be an available estimate")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reject actuals from a different completion boundary

When an estimate targets one boundary, such as merged, but the actual document reports another, such as source_present or released, reconciliation accepts both and derives comparable duration, wait, and cost errors even though they cover different scopes of work. This can create false misses and contaminate later calibration, so require actual2["completion_boundary"] to match the prior estimate's headline completion boundary before calculating errors.

Useful? React with 👍 / 👎.

Comment on lines +921 to +924
total_tokens = sum(known_token_samples) + sum(unpriced_token_samples)
known_basis = 0 if total_tokens == 0 else sum(known_token_samples) * 10_000 // total_tokens
unpriced_basis = 0 if not token_rows else 10_000 - known_basis
known_cost_q = {"p50": None, "p80": None, "p95": None} if not routes or not token_rows or sum(known_token_samples) == 0 else _q(cost_samples)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Treat fully priced zero token usage as a known zero cost

When a valid published token prior has all-zero quantiles and the request has routes with unambiguous prices, total_tokens == 0 forces coverage to 0, unpriced coverage to 10000, and every cost quantile to null. Zero modeled usage is fully priced at zero rather than unpriced, so this case should return 10000 known basis points and zero cost while continuing to reserve the unavailable/null state for missing token priors or routes.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants