We score the browser.
Never the person using it.
Every input to the score is an action a browser took — a blocked paste, a window that lost focus. Never an inference about a face, a voice or a writing style. Each observation carries its most likely innocent explanation, and nothing here decides. A person does.
No candidate is ever automatically rejected. That is enforced by a database constraint, not a code-review convention — a rule that lives in a developer's memory is not a rule.
- window lost focusthey opened the job description again
- paste into answera name they had copied from the brief
- second screen presenta docked laptop, connected all session
- network changeda phone moved between wifi and mobile data
Blocked outright during recording
- paste · copy · cut
- context menu right-click
- devtools in-page shortcuts
The candidate is told before they try, not caught after.
Observed, with context attached
- fullscreen_exit
- tab_hidden · window_blur
- network_offline
- virtual_camera_detected
Each weighted, capped, and attributed to a named observation.
And we say so on the candidate's screen
- window switching — no browser permits blocking it
- a second device in the room
A wall you pretend to have is a wall someone walks through.
A signal without its innocent reading is an accusation with a number on it.
This is what a reviewer actually sees. Deterministic — the same interview produces the same score and the same words, every time, because that text is what a candidate can demand and what an employer would have to defend.
And the most important file in it is a list of things it won't do.
Not a roadmap. Not "coming soon". Four methods we built against, deleted, and documented in the source with the reasoning intact.
AI-generated-text detection
Published error rates exceed 50% for non-native English writers, against near-zero for native speakers. Our input would be a transcript of speech, so an accent reaches the classifier twice — through the speaker, and through the recogniser's own higher error rate on accented audio. Consent does not cure a disparate impact.
Authorship comparison between answers
Attributing five short spoken answers to one person is not reliable, and its failure mode is people who code-switch or speak a second language.
Network-origin scoring
VPNs, corporate networks and Tor are far too common among privacy-conscious people, remote employees and people in censored countries to be evidence of anything alone. Recorded for a reviewer, excluded from scoring, with a test asserting the exclusion.
Automated face-match thresholds
NIST measures materially higher false non-match rates for Black, East Asian and female faces. A threshold converts a measured demographic differential into a hiring outcome. Similarity is shown as advice beside both photographs; a named person writes the reasoning.
We audit which signals fire, not only who gets hired.
Flag rates, not just selection rates
Selection-rate analysis only sees harm once it has reached a hiring decision. Measuring how often each signal fires per self-identified category surfaces a detector misreading a disability months earlier.
Past the four-fifths rule
Flag rates of 2% and 6% become favourable rates of 98% and 94% — a ratio of 0.96 that clears the 80% rule while one group is flagged three times as often. So we compare flag rates directly. There is a test asserting exactly that case.
Body-derived signals stay gated
Blink rate, gaze, head movement and pause rhythm are weighted at almost nothing, dropped entirely if a candidate asks for another way with no reason required, and withheld from scoring until a flag-rate audit against disability has run.
Small groups are withheld
Fewer than five candidates in a group and the report shows nothing — below five it is statistically meaningless and re-identifying at once. A flagged disparity opens an investigation that must close with findings and remediation.
Rights that work without an email to support.
Priced per evaluation, so cost tracks hiring volume.
No per-seat minimum. A 14-day trial, then an annual commitment.
Starter
Billed annually at $2,988
- 150 evaluations included
- 3 seats
- Push scores to your ATS
- Sandbox
Growth
Billed annually at $8,988
- 500 evaluations included
- Unlimited seats
- AI scoring of interview answers
- Identity verification
- Keel workforce planning
Enterprise
Quoted to your hiring volume
- SSO and SCIM
- Adverse-impact reporting
- ATS import and sync
- Offer workflow
You are not charged for a candidate who declines at the consent gate, or one whose interview never starts. Billing is idempotent per candidate, so an internal retry cannot double-charge you, and the usage ledger outlives deletion of the candidate record so an invoice stays reconcilable after their data is gone.
What we can prove today — and what is still in progress.
Stated plainly, because a claim you cannot check is a claim you should not make.
The questions people actually ask.
Isn't this just proctoring software?
Proctoring watches the person. Every input to our score is an action a browser took — a paste attempt, a devtools shortcut, leaving full screen. The output is an advisory band that routes to a human rather than a pass or a fail. The signals we do derive from a candidate's body are weighted near zero and excluded on request.
Can a candidate just use a phone next to their laptop?
Yes. We block what a browser can block and record what it can observe; a second device is outside both. We say this on the candidate's screen rather than implying a wall that is not there.
What if a signal fires on someone unfairly?
Every observation is logged with its most likely innocent explanation and routed to a person, never acted on automatically. A candidate can request another way with no reason required, which excludes every body-derived signal from their scoring while the observations stay recorded as evidence for the audit. The flag-rate report exists to catch a signal that misfires systematically, before anyone notices case by case.
How accurate is the fraud detection?
We do not publish an accuracy figure, because nothing here classifies a person. It counts browser events deterministically. The detectors where accuracy would be the interesting question are the four we did not ship.
Put it on the record.
Create an account and run your first interview today. Priced per evaluation, billed as the work happens. Or talk to us if you would rather be walked through it.