Methodology & Limitations

The Hoeflin Power Test Exposure Watch uses a structured taxonomy to classify public exposures of test materials. This document outlines our definitions, confidence intervals, and the limitations of our tracking approach.

Classification Taxonomy

Prompt Exposure

The exact item text, diagram, or question has been published verbatim. This indicates that the raw material of the test is publicly available, but does not necessarily imply that solutions or methods are attached.

Example: An anonymous forum post containing "Item 14: [verbatim puzzle text] — anyone know how to solve this?"

Method / Hint Exposure

A solving technique, specific hint, or partial framing that meaningfully guides the solution has been published. The actual final answer may not be present, but the cognitive leap required to solve it has been compromised.

Example: A discussion explaining that a particular numerical sequence relates to prime gaps, without giving the next number.

Candidate Answer Exposure

A proposed answer or set of candidate answers has been published. These answers are not necessarily correct or verified, but their presence online pollutes the independence of future submissions if candidates search for prior art.

Example: A pastebin file listing guessed answers for items 1-20, containing a mix of correct and incorrect responses.

Verified / Strong Solution Exposure

A correct solution or strongly verified answer has been circulated. This represents a severe compromise of the item's validity for psychometric purposes.

Example: A detailed blog post walking through the exact logic and arriving at the definitively correct answer for a high-ceiling item.

Confidence Levels

Every finding is assessed with a confidence level regarding the accuracy of the exposure claim and the persistence of the source material.

  • Low:Source is highly transient, ambiguous, or unverifiable (e.g., deleted tweet with no archive).
  • Medium:Source exists but the relevance to the specific item is debatable, or the candidate answer is clearly derived from a flawed method.
  • High:Clear, unambiguous exposure of the material on a persistent platform.
  • Verified:Confirmed by multiple researchers or matched definitively against the canonical key.

Limitations

This observatory can only track public and discoverable exposures. We recognize several inherent limitations to this approach:

  • Private Sharing: Exposures occurring in private Discord servers, direct messages, or closed study groups cannot be tracked.
  • Language Barriers: Our automated and manual monitoring primarily focuses on English-language communities. Exposures in other languages may be missed.
  • Ephemeral Content: Exposures on platforms with disappearing content (e.g., Snapchat, Instagram stories) are exceedingly difficult to archive unless caught immediately.
  • False Positives: High-range tests often employ archetypal puzzle formats. A discovered puzzle might resemble a Hoeflin item without actually being a direct leak. We use the "Changes Classification" flag to indicate when a finding is severe enough to warrant re-evaluating an item's status.