Acentecom hiring · design preview · v3.3 — calibrated against 11 already-judged people

Nine questions they can't aim at

No options to read, so there is no menu of company values to pick from. Each question hides what it measures — they think Q1 is the standard opener and Q3 is a skills test. Then an AI follow-up calls the bluff live, and the cross-read catches what one convincing answer can hide.

9questions
14minutes
0multiple choice
Seat under test Only Q3 and Q6 change — everything else is identical for every role
What the candidate reads before the first question Write like you'd talk. Your interviewer will have your answers in front of them and will ask about some of them.

Slides, not a scroll

Between the questions

One question per screen. Answer, continue, next thing arrives. Between them, a short card about the company — no question on it. It breaks the interrogation rhythm, and it sells the place to someone who has options.

    The one hard rule about these

    An interstitial must never sit immediately before a question it would prime. The intensity card in particular goes after the Sunday question, never before it — telling someone we work weekends and then asking what they did last Sunday produces a rehearsed answer and destroys the item.

    Three integrity layers

    Calling the bluff

    Open questions are more exposed to a pasted answer than multiple choice was. These three make that expensive rather than pretending it can't happen.

    01Live follow-up

    The model reads what they just wrote and asks one follow-up at its weakest point — before they can move on.

    Why this is the strongest tool available: a prepared or pasted answer has nothing underneath it, and the follow-up is generated from their own words in real time, so it cannot be anticipated. It also turns the vaguest answers — normally just a weak grade — into a second, sharper data point.

    Examples it would ask"You said you'd align with stakeholders — who specifically, and what would you ask them?" · "You mentioned a big improvement. What was the number before, and after?" · "You said you'd research it — name one source you'd actually open."
    RulesOne follow-up per question, ~30 seconds each, about 2 minutes total. Never accusatory, never tells them they're wrong. The follow-up and their reply are both graded and both shown on the card.
    02Did they write it themselves?

    Paste and keystroke telemetry, captured per answer.

    What we capturePaste events — count, size, position. One 400-character paste into an empty box is not writing. Keystrokes vs. final length — real writing makes far more keystrokes than characters; a ratio near 1.0 means the text arrived rather than was written. Time to first keystroke, typing speed and its variance (human writing is bursty, transcription is even), and tab-blur count.
    How it is usedSurfaced as a Writing band — written live / mixed / likely pasted — plus the raw numbers. It flags for the interviewer, it never auto-rejects. Someone who drafts in a notes app is not a cheat; the live follow-up is what actually resolves it.
    likely_pasted
    03The typed pledge

    Before question one, they type — by hand, paste disabled — "I will answer in my own words."

    Two reasons it earns its place. A commitment made before the task, in your own hand, measurably raises honesty on what follows — signing at the top beats signing at the bottom. And it gives us a clean typing sample from this specific person: their real cadence and error rate, which every later answer is compared against.

    Someone who types the pledge at 38 wpm with normal typos and then produces answer four at 140 wpm with none has told us something without being asked.

    The part a candidate cannot compose against

    The cross-read

    Every answer is graded on its own, and then all seven are read together. A person can write one convincing paragraph. Writing seven that stay consistent — without knowing which consistency is being checked — is a different problem. Each contradiction below becomes a quotable line for the interviewer instead of a vague bad feeling.

      The honest tradeoff: someone will paste this into a chatbot

      Open questions are more exposed to that than multiple choice was. Four things blunt it, and they compound: four of the seven demand facts only they hold — their number, their failure, their colleague, their last six months, where a chatbot can only produce a plausible blank. Generic fluency is itself the tell — the rubric scores specificity, so a polished answer with no particulars lands weak on its own merits, no detection needed. They're told the interviewer has their answers, so inflating one means defending it out loud. And paste/dwell telemetry raises a flag for the interviewer, never an automatic rejection.

      Sample result · what everyone sees — recruiters, hiring managers, Uzi, Liam

      Marta R.

      Work
      B
      Smart
      A
      Loves the job
      D
      CraftSolid — scored at L1
      DirectionArgues, then executes
      AIPractitioner — still moving
      IntegrityClean — no paste events

      Direction is a new band, added 2026-08-09 after calibration. It reads one thing: when they were told to do something they disagreed with, did they do it — and if they didn't, did anyone find out from them? It exists because the run passed a real media buyer the company had removed: he graded A on smart, B on work and B on loves-the-job, and was let go for weeks of ignoring a direct one-line test request. The three pillars are correct and they are still only three — this reports alongside them and never gates, the same as Craft and AI. Bands: executes and argues after · executes reluctantly · argues then executes · decides alone.

      Recommendation Re-ask — decision deferred

      Contradictions found across the seven

      • Q1 ↔ Q4Wants to master long-form scripting (Q1), and would delete scriptwriting from the job if she could (Q4). One of those is not true.
      • Q2 ↔ Q5Condemns her worst colleague for blaming other teams (Q2), then attributes her own missed delivery entirely to a downstream team (Q5).

      Flags

      no_numberinconsistent_across_answerspossible_comprehension_issue

      Three flags were added 2026-08-09 after calibration: wont_take_direction unilateral_on_process output_without_outcome. The last one closes a plain gap — producing volume with nothing to show for it is the single most common removal cause in the meeting record, and it had no flag. no_number fires when a candidate cannot produce a figure; it does not fire when they produce the wrong kind of figure, which is what "29 creatives in 22 days" is.

      Ask her in the call

      1. You said you want to get great at long-form scripting, and also that you'd delete scriptwriting. Walk me through that.
      2. The project you're proudest of — what was the number, and which part was you?
      3. Tell me about the delivery that slipped. Who knew, and when?
      What the candidate sees at the end: "Thanks — this is with the hiring team." No score, no band, no feedback, ever. The rubric never reaches their browser and grading happens server-side after submission, so there is nothing readable in the network tab either. A test whose key leaks is dead.

      Preview only — nothing is deployed and nothing is calibrated yet. Best next step: give me two candidates you'd hire and two you wouldn't, and I'll run their answers past this rubric to see whether it separates them the way you would.