mirror of
https://github.com/merlinhu1/Battle-tested-Research-Prompts.git
synced 2026-08-26 12:20:55 +02:00
- Anthropic Automated W2S Researcher: PGR 0.97 vs 0.23 human baseline; full prompt stack published (system prompt, skill, critic prompts) - SPARK pathology agents: Nature Medicine, survival-validated biomarkers; full CrewAI agents/tasks prompts published - OpenAI graviton amplitudes: rejected - no prompt ever released
185 B
185 B
Question: {{ question }}
Proposed Answer: {{ choice }}
Critique: {{ critique }}
Based on the question, proposed answer, and critique, is the answer correct? Answer (True or False):