🧪

Copilot Playground

Sign in with your Shortlister account to continue.

Copilot Playground

Takes ~60–90 seconds

Real Copilot. Real rigor.

Every answer is evaluated against the same scorecard, every time.

Simulate an interview

See Copilot in action in seconds.

Tell us what to simulate

Choose a role / programme. We'll create the interview and participants.

Watch Copilot work

See Copilot evaluate every answer using a consistent scorecard.

Explore the results

Review outcomes, feedback and the evidence behind every score.

Build a question scorecard (Stage 1 → Stage 2)

Used for calibration

Saved questions

LabelHubStatusCreated

Build an interview

A named, ordered set of questions. Build it once, then every participant who takes it gets exactly this question set for Wash-Up.

Saved interviews

NameQuestionsFluency thresholdSimulationCreated

New participant

Quality only matters if you use "Generate Transcript" below — it has no effect if you paste in your own transcript.

Or bulk-upload via CSV

Creates multiple participants at once — with their own real transcripts, already filled in — against the interview selected above. Max 40 rows per upload. Download the template first: it has the exact columns for that interview's current questions — each question gets two columns, its transcript followed by an optional "Question - Human Score" column, if you already have a human reviewer's score to hand for that answer. Files with more than 40 rows, or whose columns don't match the template exactly, are rejected automatically before anything is created. This does not auto-run Evaluate or Wash-Up — those stay manual steps per participant, same as any other participant.

Auto-generate test participants

Creates fictitious participants against an interview — generates a name, a synthetic transcript for every question, then automatically runs Evaluate and Wash-Up, so each one comes out fully scored. This makes real API calls (Stage 5, 6, 3 x questions, and 4 per participant), so it isn't free/instant unless Offline test mode is on.

Participants

NameInterviewGenderQualityFluencyWash-UpCreated

Evaluate a response (Stage 3)

"Evaluate All" runs Stage 3 for every pending or previously-failed answer that already has a transcript, across every participant and every question — the CSV upload template fills transcripts in but doesn't evaluate them, so this saves triggering each one by hand. Answers already evaluated are left as they are. "Re-score out of date" appears only when a question's scorecard has been rebuilt since some answers were scored: it re-scores exactly those, then recomputes the interview outcomes that depended on them. Human review scores are never touched by either.

Interview Wash-Up (Stage 4)

"Run All Wash-Ups" runs Stage 4 for every participant whose questions are all evaluated (or disabled) — the same completeness check "Run Wash-Up" enforces one participant at a time below. Anyone with a pending or failed evaluation is skipped, not treated as an error; a participant already washed up is simply re-run against their latest evaluations.

Wash-Up Summary

A grid of every participant who has completed Wash-Up for a chosen interview, one column per question.

Human-Copilot Calibration

Compares a human reviewer's score against Copilot's own level for the same answer, one question at a time (Stage 8) — for every washed-up participant in the interview below who has both a human score (set on the Evaluate Answers tab) and a max score set on the question (Build Interview tab). Any pair missing either, or not yet evaluated, is skipped automatically rather than shown as an error. Nothing is calibrated until you click Run Calibration. Re-running only does the work that's needed: pairs already calibrated whose answer, score or question hasn't changed since are reused, so a run after adding one participant calibrates that participant. Use Re-run everything to rebuild the interview's results from scratch.

OpenAI connection

Applies to every scored stage. If your model rejects it, calls fall back to standard speed rather than failing — use Test Connection to check.

Logo

Shown in the header, to the left of "Copilot Playground". SVG or PNG only.

No logo uploaded yet.

Sidebar "About" text

Optional. Shown as plain text below the main navigation in the left sidebar — a good place for a one- or two-line note about what this tool is or who it's for. Leave blank to hide it.

Institution name

Who is making the offer, in the words a participant would recognise (e.g. "Shortlister University"). Used by the Personal Offer Story to establish what the participant was being considered for — combined with the interview name, so "Shortlister University" + "BA Primary Teaching". Leave blank and the story names only the interview.

Simulation roles / programmes

The "Ready-made" list in Simulation Mode's Role / programme dropdown. Each role asks its questions in order — a run takes the first N of them, so question 1 is the one every run shows. Copilot still builds every scorecard itself; this only decides what gets asked.

Text-to-Speech ("Speak the transcript")

Applies to every "Speak the transcript" button across the app. Changing these doesn't affect audio already generated and cached — the next click regenerates using whatever's set here.

Advanced audio settings

Voice depends on each participant's gender (set when the participant is created) — not one voice for everyone.

Default mode

Which experience the app opens in. Set this to Simulation Mode for a client-facing install; leave it on Hands-on Mode if you mostly work in the detailed Playground yourself. Either way both modes stay one click apart in the header.

UI Copy

The major page titles, panel headings and explanatory text shown throughout the app, in one editable block of JSON. Copy it out to any text editor, make changes, and paste it back in — anything you don't override falls back to the app's built-in default wording.

Reset Participant Datawipes every participant, evaluation, Wash-Up result and Calibration result, and clears cached "Speak the transcript" audio. Your questions, interviews, API key, model setting, TTS settings, logo, About text, and prompt version history are all kept as-is — use this to clear out test participants between rounds without rebuilding your question bank. This can't be undone.
Reset Everythingwipes every question, interview, participant, evaluation, Wash-Up result and Calibration result, and clears cached "Speak the transcript" audio. Your API key, model setting, TTS settings, logo, About text, and prompt version history are all kept as-is. This can't be undone.