Files
codex-game-studio/.agents/skills/cgs-test-evidence-review/SKILL.md
T
7e8846a7fb feat: add deterministic model tier routing (#14)
* feat: add deterministic model tier routing

Define Sol, Terra, and Luna policies with fixed fallbacks, explicit surface assignments, runtime precedence, and Luna escalation. Validate generated metadata and expose tier overrides across role, task, and orchestration commands.

* fix: infer model tiers from surface models

* fix: rebalance prompt model routing

* fix: decouple prompt effort routing

---------

Co-authored-by: MerlinH <merlinh221@gmail.com>
2026-07-10 18:49:20 +10:00

2.6 KiB

name, description, model, model_reasoning_effort, argument-hint, primary-agent, tool-policy, isolation, source-reference, source-hash, user-invocable
name description model model_reasoning_effort argument-hint primary-agent tool-policy isolation source-reference source-hash user-invocable
cgs-test-evidence-review Use for test evidence review tasks that review validation output, logs, screenshots, recordings, and manual checks for sufficiency and gaps; produce verification evidence, changed or proposed files, and handoff boundaries. gpt-5.6-luna medium Describe the test-evidence-review objective, target files/assets, constraints, and verification evidence. qa-playtester read/edit/shell/tests/git as needed within the repository write policy repository-root; no init-time generated prompt bodies .claude/skills/test-evidence-review/SKILL.md 04cdf0b7e94ab2dab58850bf1a70a2a6f33810ea7abb6102eaec8f2d00b9026e true

Codex Game Studio Test Evidence Review

Use this skill for test evidence review work in Template Game.

Objective

Review validation output, logs, screenshots, recordings, and manual checks for sufficiency and gaps.

Inputs

  • AGENTS.md
  • .codex/studio.json
  • task-relevant files named by the user or task record
  • design/gdd.md
  • production/session-state/active.md
  • tests/

Arguments

  • Objective or user request.
  • Target files, scenes, assets, or docs.
  • Constraints, deadlines, acceptance criteria, and verification command when known.

Procedure

  1. Clarify the requested review validation output, logs, screenshots, recordings, and manual checks for sufficiency and gaps. and identify the current project stage.
  2. Collect evidence for Evidence, Gap, Confidence.
  3. Run or define the focused validation loop before reporting conclusions.
  4. Produce the requested artifact or review with clear file paths and verification evidence.

Write Targets

  • tests/
  • production/session-state/
  • prototypes/

Output Contract

  • Summary
  • Evidence
  • Gap
  • Confidence
  • Follow-up
  • Risks
  • Changed files or proposed files
  • Verification evidence
  • Next owner or decision

Quality Gates

  • Evidence
  • Gap
  • Confidence
  • Follow-up
  • Scope remains bounded to the current task and project stage.
  • Report labels unverified assumptions separately from evidence.

Decision Gates

  • Continue only when the expected output can be verified or clearly labeled as a plan.
  • Escalate to producer or qa-playtester when scope, ownership, or acceptance evidence is ambiguous.
  • Stop before broad rewrites, generated prompt mirrors, or hidden lifecycle behavior.

Handoff

Report changed files, verification evidence, remaining risks, and the next owner or decision. Do not imply hidden hooks or autonomous background work.