mirror of
https://github.com/merlinhu1/Battle-tested-Research-Prompts.git
synced 2026-08-26 12:20:55 +02:00
- Anthropic Automated W2S Researcher: PGR 0.97 vs 0.23 human baseline; full prompt stack published (system prompt, skill, critic prompts) - SPARK pathology agents: Nature Medicine, survival-validated biomarkers; full CrewAI agents/tasks prompts published - OpenAI graviton amplitudes: rejected - no prompt ever released
231 B
231 B
Question: {{ question }}
Option A: {{ option_a }} Critique of A: {{ critique_a }}
Option B: {{ option_b }} Critique of B: {{ critique_b }}
Based on the question, options, and critiques, which option is better? Answer (A or B):