Sign in with your Shortlister account to continue.
Copilot Playground
Takes ~60–90 seconds
Real Copilot. Real rigor.
Every answer is evaluated against the same scorecard, every time.
Simulate an interview
See Copilot in action in seconds.
Tell us what to simulate
Choose a role / programme. We'll create the interview and participants.
Watch Copilot work
See Copilot evaluate every answer using a consistent scorecard.
Explore the results
Review outcomes, feedback and the evidence behind every score.
Open a completed simulation to explore its results again — nothing is re-run.
Simulation in progress...
Sit back and watch Copilot evaluate responses in real time.
Creating interview
Generating participants
Evaluating answers
Evaluating interview
The results are ready
Every answer has been evaluated. Open the results when you're ready.
Participants
One square per answer, filling in as Copilot evaluates it. The recommendation on the right is drawn from all of them together.
The scorecard Copilot will apply
Every answer is measured against these same requirements — this is what makes the assessment consistent.
Explore the results
Review participant outcomes and dive into the detail.
Simulated interview
Answers that failed to evaluate
Participant outcomes
Click a participant to see their detailed results.
Interview questions
We're delighted to be offering you a place on our course.
Build a question scorecard (Stage 1 → Stage 2)
Editing a saved question — Build & Save will update it in place. Click "New Question" to start a blank one instead.
Locked - delete & recreate question to amend
Can still be changed — no answers evaluated yet
Used for calibration
Label, max score and evaluation status — no rebuild
Building this question — the form is locked until it finishes so the result can't land on a different question.
Saved questions
Label
Hub
Status
Created
Build an interview
A named, ordered set of questions. Build it once, then every participant who takes
it gets exactly this question set for Wash-Up.
Saved interviews
Name
Questions
Fluency threshold
Simulation
Created
New participant
Quality only matters if you use "Generate Transcript" below —
it has no effect if you paste in your own transcript.
Or bulk-upload via CSV
Creates multiple participants at once — with their own real transcripts, already filled
in — against the interview selected above. Max 40 rows per upload. Download the template
first: it has the exact columns for that interview's current questions — each question gets two
columns, its transcript followed by an optional "Question - Human Score" column, if you already
have a human reviewer's score to hand for that answer. Files with more than 40 rows, or whose columns
don't match the template exactly, are rejected automatically before anything is created. This does
not auto-run Evaluate or Wash-Up — those stay manual steps per participant, same as any other
participant.
Auto-generate test participants
Creates fictitious participants against an interview — generates a name, a synthetic
transcript for every question, then automatically runs Evaluate and Wash-Up, so each one comes out
fully scored. This makes real API calls (Stage 5, 6, 3 x questions, and 4 per participant), so it isn't
free/instant unless Offline test mode is on.
Participants
Name
Interview
Gender
Quality
Fluency
Wash-Up
Created
Evaluate a response (Stage 3)
"Evaluate All" runs Stage 3 for every pending or previously-failed answer
that already has a transcript, across every participant and every question — the CSV
upload template fills transcripts in but doesn't evaluate them, so this saves triggering
each one by hand. Answers already evaluated are left as they are. "Re-score out of date"
appears only when a question's scorecard has been rebuilt since some answers were
scored: it re-scores exactly those, then recomputes the interview outcomes that depended
on them. Human review scores are never touched by either.
Interview Wash-Up (Stage 4)
"Run All Wash-Ups" runs Stage 4 for every participant whose questions are all
evaluated (or disabled) — the same completeness check "Run Wash-Up" enforces one participant
at a time below. Anyone with a pending or failed evaluation is skipped, not treated as an
error; a participant already washed up is simply re-run against their latest evaluations.
Wash-Up Summary
A grid of every participant who has completed Wash-Up for a chosen interview, one column per question.
Human-Copilot Calibration
Compares a human reviewer's score against Copilot's own level for the same answer, one
question at a time (Stage 8) — for every washed-up participant in the interview below who has both a
human score (set on the Evaluate Answers tab) and a max score set on the question (Build Interview
tab). Any pair missing either, or not yet evaluated, is skipped automatically rather than shown as an
error. Nothing is calibrated until you click Run Calibration. Re-running only does the work that's
needed: pairs already calibrated whose answer, score or question hasn't changed since are reused, so
a run after adding one participant calibrates that participant. Use Re-run everything to rebuild the
interview's results from scratch.
Participants that failed to calibrate
Interview-level alignment
The human's per-question scores, put through the same signal method Copilot
uses for the interview as a whole, and compared with Copilot's own recommendation. A participant
needs a human score on every question Copilot judged to be comparable — this is a translation of
the human's scores, not a second opinion on them, so nothing here is estimated around a gap.
Question-level alignment
Edit wording freely. Do not change the field names in a prompt's Input/Output schema — the
other 3 stages depend on those exact names to hand data between each other.
OpenAI connection
Applies to every scored stage. If your model rejects it, calls fall back to standard speed rather than failing — use Test Connection to check.
Logo
Shown in the header, to the left of "Copilot Playground". SVG or PNG only.
No logo uploaded yet.
Sidebar "About" text
Optional. Shown as plain text below the main navigation in the left
sidebar — a good place for a one- or two-line note about what this tool is or who
it's for. Leave blank to hide it.
Institution name
Who is making the offer, in the words a participant would recognise (e.g.
"Shortlister University"). Used by the Personal Offer Story to establish what the
participant was being considered for — combined with the interview name, so
"Shortlister University" + "BA Primary Teaching". Leave blank and the story names
only the interview.
Simulation roles / programmes
The "Ready-made" list in Simulation Mode's Role / programme dropdown. Each role
asks its questions in order — a run takes the first N of them, so question 1 is the one
every run shows. Copilot still builds every scorecard itself; this only decides what
gets asked.
Text-to-Speech ("Speak the transcript")
Applies to every "Speak the transcript" button across the app. Changing these doesn't
affect audio already generated and cached — the next click regenerates using whatever's set here.
Advanced audio settings
Voice depends on each participant's gender (set when the
participant is created) — not one voice for everyone.
Default mode
Which experience the app opens in. Set this to Simulation Mode for a
client-facing install; leave it on Hands-on Mode if you mostly work in the detailed Playground
yourself. Either way both modes stay one click apart in the header.
UI Copy
The major page titles, panel headings and explanatory text shown throughout the app, in
one editable block of JSON. Copy it out to any text editor, make changes, and paste it back
in — anything you don't override falls back to the app's built-in default wording.
Reset Participant Data — wipes every participant, evaluation, Wash-Up result and
Calibration result, and clears cached "Speak the transcript" audio. Your questions,
interviews, API key, model setting, TTS settings, logo, About text, and prompt version
history are all kept as-is — use this to clear out test participants between rounds
without rebuilding your question bank. This can't be undone.
Reset Everything — wipes every question, interview, participant, evaluation,
Wash-Up result and Calibration result, and clears cached "Speak the transcript" audio.
Your API key, model setting, TTS settings, logo, About text, and prompt version history
are all kept as-is. This can't be undone.