Evaluation methodology
How We Evaluate AI Philosopher Debates
PhiloMind AI uses a versioned evaluation process before treating a philosopher as debate-ready. The current v1 method combines pairwise philosophical cases, strict automated integrity checks, and a separate human quality gate. This page documents the method and its limits; it does not claim that an AI interpretation is the historical philosopher or that every possible response is error-free.
The v1 methodology at a glance
The current Debate evaluation dataset covers five debate personas: Socrates, Plato, Epictetus, Friedrich Nietzsche, and Immanuel Kant. It contains ten curated cases, which gives one case for every possible pair among those five philosophers.
A case is not considered ready because a model merely produces fluent text. All critical automated checks must pass, and a human reviewer must separately approve the quality of the philosophical exchange. The current human threshold requires every reviewed dimension to score at least 4 out of 5 and the average across required dimensions to reach at least 4.2 out of 5.
- 5 debate personas in the current v1 evaluation set
- 10 curated pairwise cases: one for every philosopher pair
- Zero tolerance for the defined automated critical failures
- Human review required after automated checks
- Minimum 4/5 on every required human dimension and 4.2/5 average
What one evaluation case tests
Each case begins with a concrete philosophical disagreement rather than a biography question. The standard structure expects both philosophers to participate across two autonomous rounds, then uses a targeted user intervention to put pressure on the exchange.
The intervention is deliberately specific. It may ask for a concrete case, require a counterexample, force one philosopher to answer the strongest point in the opponent's position, or test whether a philosopher can disagree without caricaturing the other side. This makes the evaluation about interaction quality rather than isolated one-shot answers.
- A named philosophical disagreement with two defensible perspectives
- Expected turns from both participants across the debate structure
- A targeted intervention that tests direct engagement under pressure
- Opponent-aware responses rather than two unrelated monologues
Automated critical integrity checks
The automated layer is designed to catch failures that should block readiness even when the prose sounds persuasive. These checks focus on structure, grounding, philosopher isolation, and leakage rather than trying to assign a machine score to philosophical truth.
- The required logical turns exist and the evaluation run completes
- Only the expected speaker and opponent appear in each turn
- Curated benign cases produce substantive debate rather than an inappropriate safety or insufficient-evidence fallback
- The response engages and naturally identifies its opponent
- The AI does not claim to have won the debate or present itself as the literal historical philosopher
- Private evidence identifiers, internal runtime instructions, and planning fragments do not leak into visible prose
- The same long response is not simply repeated by a philosopher within a debate
- Grounded or interpretive turns include citations
- Citation ownership stays with the philosopher who is currently speaking rather than crossing between personas
The human quality gate
Automated integrity is necessary but not sufficient. A response can satisfy formatting and citation rules while still being philosophically weak, generic, or unfair to the opponent. For that reason the v1 readiness policy requires an explicit human review before a case can pass.
The reviewer scores three dimensions: direct engagement, persona distinctiveness, and fair opponent representation. A generated report that has not been human reviewed remains incomplete for readiness purposes, regardless of how many automated checks it passes.
- Direct engagement: does the response actually answer the opponent's argument?
- Persona distinctiveness: does the reasoning remain recognizably specific to this philosopher rather than generic debate prose?
- Fair opponent representation: does the response challenge the other position without replacing it with an easier caricature?
The current pairwise case matrix
The v1 dataset covers the complete pairwise matrix for the five currently evaluated Debate personas. The topics are chosen to create genuine pressure between different methods or philosophical commitments.
- Socrates vs Epictetus — control, judgment, examination, and the good life
- Socrates vs Immanuel Kant — definitions, universal duties, and moral terms
- Socrates vs Friedrich Nietzsche — examination, genealogy, and inherited values
- Socrates vs Plato — political justice, inquiry, and knowledge of the good
- Epictetus vs Immanuel Kant — disciplined choice, duty, and uncontrollable outcomes
- Epictetus vs Friedrich Nietzsche — adversity, judgment, and self-overcoming
- Epictetus vs Plato — inner freedom, education of the soul, and political responsibility
- Immanuel Kant vs Friedrich Nietzsche — universal morality and historically situated valuation
- Immanuel Kant vs Plato — rational autonomy and an objective Good
- Friedrich Nietzsche vs Plato — discovered truth, higher values, and revaluation
What source grounding means in this evaluation
The Debate evaluator checks that grounded or interpretive turns carry citations and that those citations stay isolated to the philosopher who is speaking. This is important because a plausible sentence backed by another philosopher's evidence would still be a grounding failure.
Citation presence is not the same thing as proving every philosophical claim correct. Separate source and release gates are used for questions such as exact-quotation integrity, source attribution, witness or edition accuracy, and broader philosopher-specific acceptance. The pairwise Debate methodology should therefore be read as one layer in a larger quality system, not as a universal historical-authenticity score.
What a passing evaluation does not mean
A pass means that a particular version met the defined readiness gates on the curated cases that were actually tested. It does not mean the system has recreated a historical person, settled a philosophical disagreement, or proven correctness on every future prompt.
Evaluation evidence is version-sensitive. Changes to a persona, retrieval behavior, source set, model behavior, renderer, or debate orchestration can change the relevant risk surface. Older evidence should not be treated as permanent proof for materially different software.
- Not a claim of literal historical identity
- Not a guarantee for every possible user question
- Not a single score for philosophical truth
- Not permanent evidence after material system changes
How we will publish evaluation results
PhiloMind AI will not turn an automatically generated but unreviewed report into a public performance claim. Public result pages should identify the evaluation version, tested scope, human-review status, limitations, and the difference between integrity checks and subjective quality judgments.
We also keep private prompt material, internal evidence identifiers, credentials, and security-sensitive runtime details out of public evaluation artifacts. The goal is enough transparency to understand and scrutinize the method without publishing information that would weaken the system's security or evaluation integrity.
Source transparency
Every conversation on PhiloMind AI is AI-generated and shaped by approved philosophical sources, documented reasoning methods, and explicit source disclosures when available. It is not the historical person, and it may contain uncertainty or limitations.
