Methodology & Limitations
The Hoeflin Power Test Exposure Watch uses a structured taxonomy to classify public exposures of test materials. This document outlines our definitions, confidence intervals, and the limitations of our tracking approach.
Classification Taxonomy
Prompt Exposure
The exact item text, diagram, or question has been published verbatim. This indicates that the raw material of the test is publicly available, but does not necessarily imply that solutions or methods are attached.
Method / Hint Exposure
A solving technique, specific hint, or partial framing that meaningfully guides the solution has been published. The actual final answer may not be present, but the cognitive leap required to solve it has been compromised.
Candidate Answer Exposure
A proposed answer or set of candidate answers has been published. These answers are not necessarily correct or verified, but their presence online pollutes the independence of future submissions if candidates search for prior art.
Verified / Strong Solution Exposure
A correct solution or strongly verified answer has been circulated. This represents a severe compromise of the item's validity for psychometric purposes.
Confidence Levels
Every finding is assessed with a confidence level regarding the accuracy of the exposure claim and the persistence of the source material.
- Low:Source is highly transient, ambiguous, or unverifiable (e.g., deleted tweet with no archive).
- Medium:Source exists but the relevance to the specific item is debatable, or the candidate answer is clearly derived from a flawed method.
- High:Clear, unambiguous exposure of the material on a persistent platform.
- Verified:Confirmed by multiple researchers or matched definitively against the canonical key.
Limitations
This observatory can only track public and discoverable exposures. We recognize several inherent limitations to this approach:
- Private Sharing: Exposures occurring in private Discord servers, direct messages, or closed study groups cannot be tracked.
- Language Barriers: Our automated and manual monitoring primarily focuses on English-language communities. Exposures in other languages may be missed.
- Ephemeral Content: Exposures on platforms with disappearing content (e.g., Snapchat, Instagram stories) are exceedingly difficult to archive unless caught immediately.
- False Positives: High-range tests often employ archetypal puzzle formats. A discovered puzzle might resemble a Hoeflin item without actually being a direct leak. We use the "Changes Classification" flag to indicate when a finding is severe enough to warrant re-evaluating an item's status.