Introducing Council 1.2: Anonymous External AI Answer Review
The latest iteration of the Council AI evaluation tool, version 1.2, introduces a significant new capability: the ability to anonymously integrate external AI-generated answers into its multi-model blind review process. This update allows users to paste an answer from any source—be it ChatGPT, Gemini, Bard, Claude, a colleague's response, or even a draft from a human expert—and have it evaluated alongside answers from the models directly integrated into Council. The core principle remains: every model critiques every other model's output without knowing its origin, ensuring a fair and unbiased assessment of quality and accuracy.
Council's unique approach is to present a single question to multiple AI models simultaneously. Once responses are generated, the system anonymizes them. Each participating model is then tasked with reviewing and critiquing the anonymized answers from its peers. This process strips away brand prestige or perceived model authority, forcing each AI to evaluate content based on its merits. A key output of this system is a quantitative score, typically on a 0-100 scale, indicating how closely the models' assessments align. It also highlights instances where one model's answer diverged significantly from the consensus, potentially identifying unique insights or critical errors.
The addition of the 'guest seat' in version 1.2 is particularly noteworthy. Previously, Council focused solely on comparing the native capabilities of the models it could directly access. Now, users can inject any arbitrary answer into the evaluation fray. This transforms Council from a comparative benchmark of specific models into a more versatile quality assurance and analysis tool. Imagine a scenario where a company is developing its own internal AI assistant. They could use Council 1.2 to anonymously pit their in-house model's answers against those of leading commercial models, or even against curated expert responses, all while ensuring the models providing the critiques don't know which answer belongs to whom.
The Mechanics of Blind Critiques
The review process in Council 1.2 is designed to simulate a peer-review committee where identities are concealed. When a user submits a prompt, Council dispatches it to a selected group of AI models. After receiving the responses, the system anonymizes them, often by assigning generic labels like 'Model A', 'Model B', etc. These anonymized answers are then fed back into the same models, along with the original prompt, and each model is instructed to evaluate them. The instructions typically involve assessing accuracy, relevance, coherence, completeness, and adherence to any specified constraints in the prompt.
The output of this process is multifaceted. Users receive a score indicating the degree of agreement among the critiquing models. This score is crucial for understanding the consensus on a given answer. For instance, if all models agree that an answer is excellent, the score will be high. Conversely, if models offer wildly different assessments—one praising an answer while another condemns it—the score will be low, signaling a contentious or potentially problematic response. Furthermore, the system identifies which specific model, if any, provided an outlier opinion. This could be an AI that uniquely recognized a subtle flaw or, conversely, an AI that missed a crucial aspect everyone else caught.
The 'guest seat' feature operates on the same anonymized critique principle. When a user pastes an external answer, it is assigned a temporary, anonymous identifier. This answer then enters the pool of responses to be critiqued by the other models. The models providing the critiques do not know if the guest answer came from a competitor, a human, or even one of the models they are already evaluating under a different label. This ensures that the guest answer receives an unbiased evaluation, free from any preconceived notions about its origin. The resulting critiques and scores for the guest answer are then presented to the user, offering a valuable data point for comparison.
Why This Matters: Beyond Benchmarking
Council 1.2 moves beyond simple performance benchmarking. While comparing raw output scores between models is useful, understanding *why* models agree or disagree is far more insightful. The blind critique mechanism provides a window into the models' reasoning processes, albeit indirectly. When models consistently flag a particular answer as weak due to factual inaccuracies, it suggests a shared understanding of verifiable truth within that model cohort. When they disagree, it might point to differing training data, subjective interpretations, or varying levels of sophistication in handling nuance.
The integration of external answers elevates Council's utility for practical applications. For developers building AI-powered features, this means they can test their prototype responses against established models without revealing their hand. This is akin to a chef anonymously submitting their new dish to a panel of renowned critics who are also tasting each other's signature dishes. The feedback is pure, unadulterated critique. For researchers, it offers a method to probe the comparative strengths and weaknesses of different AI architectures or fine-tuning approaches by using Council as a consistent evaluation framework.
Consider a scenario where a legal team is drafting an AI-assisted document. They could input the draft answer into Council 1.2 alongside answers from models like GPT-4 and Claude 3. The other models would then critique the draft, along with their own outputs. If the draft receives consistently low scores or specific, shared criticisms from the other AIs, it signals a need for revision. This process is far more granular than a simple side-by-side comparison; it's an active interrogation of the answer's quality by multiple AI intelligences.
The surprise element here is not the concept of blind reviews, which is common in academic and scientific circles. The genuine surprise is the application of this rigorous methodology to the rapidly evolving, often opaque, world of large language models, and the extension of it to arbitrary, user-supplied answers. It democratizes a level of AI evaluation that was previously difficult to achieve without specialized infrastructure or proprietary access to model review capabilities. This makes sophisticated AI quality assessment accessible to a broader audience, from individual developers to enterprise teams.
What remains to be seen is how effectively Council 1.2 can scale its critique capabilities as the number of integrated models and the complexity of user prompts increase. The computational cost of having every model critique every other model, plus an external answer, could become substantial. Furthermore, the nuances of AI critique itself—whether an AI's critique is truly objective or subtly biased by its own architecture—will continue to be a subject of research and refinement.