Last updated June 5, 2026.
Cognitive Testing Methodology
Cognitive Testing builds public analogs of serious cognitive ability tests. We study the structure of established professional instruments, identify the reasoning demands they are meant to measure, and create original online item sets that approximate those task families without copying proprietary content.
Our goal is not to replace a licensed, supervised psychological evaluation. The goal is to make careful, transparent, statistically reviewed cognitive testing available online while keeping the limits of online administration visible.
How We Build Tests
Each test begins with a professional-test review. We examine the kinds of tasks used in respected intelligence batteries and group tests: matrix reasoning, verbal abstraction, numerical induction, spatial relations, working memory style tasks, speeded mixed reasoning, and other broad cognitive formats.
From that review, we build public analogs. An analog is an original test form that targets the same broad cognitive operation as a professional test section, but uses new prompts, new stimuli, new answer choices, and new scoring data. We do not copy protected items or reproduce proprietary forms.
Two-Stage Norming
We administer tests in two complementary ways. First, we hold small norming events with roughly 50 to 100 participants. These sessions give us a controlled first look at timing, item difficulty, completion behavior, score spread, and obvious item problems before relying on broader internet data.
Second, we collect online data using cheat-resistant filtering. Online data is valuable because it lets us evaluate items at larger scale, but it has to be cleaned aggressively. We use first-attempt rules, tab-out and focus-loss flags, repeated-attempt checks, response-completion filters, timing review, and IP or abuse-signal screening where available.
| Source | Purpose | Controls |
|---|---|---|
| Small norming events | Initial calibration with 50-100 participants. | Known administration window, direct completion monitoring, immediate item review. |
| Filtered online data | Larger-scale norm refinement and item statistics. | First attempts, tab-out exclusions, timing checks, repeat detection, and abuse-signal screening. |
Those two sources are compared rather than blindly pooled. If an online curve behaves differently from an event curve, we investigate whether the difference reflects sample composition, item exposure, careless responding, cheating, or a real scoring issue.
Anti-Cheat Data Collection
Online IQ testing has a basic problem: people can use outside help, retake tests, switch tabs, or search for answers. We cannot make an online test identical to a supervised clinical administration, so we treat clean data collection as a core part of the methodology.
- Norming data favors first completed attempts rather than repeated practice attempts.
- Attempts with tab-out, window-focus loss, abnormal timing, or incomplete response patterns can be excluded from norm generation.
- Repeated IPs and suspicious traffic sources are screened before being used for score calibration.
- Generated and rotating item formats reduce long-term answer exposure.
- Retest data is kept useful for practice-effect analysis, but it is not treated the same as first-attempt norming data.
The result is not "cheat proof." It is a dataset designed to be resistant to the most common forms of online contamination, with exclusions applied before the data is used to shape scoring tables.
Professional Review
Items and data are reviewed by a licensed psychologist and a psychometrician. Their review focuses on whether the items are understandable, whether they plausibly measure the intended cognitive process, whether distractors behave appropriately, and whether the observed data supports keeping the item in the active form.
Review is not a one-time rubber stamp. As additional event and online data accumulates, item behavior is checked again. Items that show poor discrimination, confusing wording, abnormal timing, or weak relationship to the total score can be revised, retired, or separated from score reporting.
Reliability And G-Loading
For each serious score, we monitor internal consistency and general-factor behavior. Cronbach's alpha is used as a reliability check: it estimates whether the items behave like a coherent measurement set rather than a pile of unrelated puzzles. We also examine g-loading, meaning how strongly the test relates to the broad general cognitive factor expected from intelligence testing.
We benchmark these statistics against professional expectations. When we say a test is in the professional range, we mean that its reliability metrics and general-factor behavior are comparable to what is expected from serious group-administered cognitive tests, not that an online score is identical to a full individually administered clinical battery.
| Metric | What it tells us | How it is used |
|---|---|---|
| Cronbach's alpha | Whether items work together consistently. | Flags noisy, redundant, or poorly aligned forms. |
| Item discrimination | Whether an item separates stronger and weaker performers. | Identifies items to keep, revise, or remove. |
| G-loading | Whether the test tracks broad cognitive ability rather than a narrow trick. | Checks that the analog behaves like a serious cognitive test. |
| Score distribution | Whether raw scores spread enough for useful interpretation. | Supports norm tables, percentile estimates, and score ranges. |
Mainstream Test Comparisons
Where possible, participants also take mainstream cognitive tests or established benchmark instruments. These comparison scores let us check whether Cognitive Testing forms behave like real intelligence measures rather than only looking good internally.
External comparison is used to evaluate convergent validity: people who perform well on respected cognitive tests should, on average, perform well on our analogs too. If a test has strong internal consistency but weak agreement with mainstream benchmarks, that is a warning sign that it may be measuring something too narrow or idiosyncratic.
These comparisons are also used to keep score claims restrained. A Cognitive Testing score is an estimate from a specific online form. It should be interpreted with the test's reliability, norm sample, confidence interval, and administration limits in mind.
What The Score Means
The reported score is based on how a person's performance compares with the cleaned norming sample for that test. We use the norming data, reliability estimates, and validation checks to convert raw performance into an interpretable score range.
A stronger score means the person performed better than a larger share of the norming sample on that form. It does not mean the person has received a clinical diagnosis, a school-placement evaluation, or a complete cognitive profile. Online testing is best treated as a serious estimate, not as the final word on a person's ability.