How to structure a 3 minute recorded MMI answer
Every answer framework you can find was built for a six minute conversation with a human who prompts you. A recorded station gives you 30 seconds, one take and no help. Here is what changes.
written by cutline. published . every fact on this page comes from a University of Auckland or Acuity Insights page, linked where it is used.
Almost every MMI answer framework you can find was written for a six to eight minute conversation with a person sitting opposite you. In that format the structure barely has to hold, because the interviewer will prompt you when you stall, follow up when you are vague, and nod when you are on track. None of that exists here.
A recorded Auckland station gives you 30 seconds to read, three minutes to talk, and one take. Nobody reacts. Nobody rescues you. The structure is the only thing holding the answer up, so it has to be one you can execute without help, from a standing start, on a scenario you have never seen.
First, the arithmetic nobody does
Three minutes is a specific amount of talking and it is worth knowing what it feels like before you meet it under pressure. At a comfortable, clearly articulated pace of roughly 140 to 150 words a minute, three minutes is somewhere around 420 to 450 spoken words.
That is our arithmetic and not a UoA figure, so treat it as a rule of thumb rather than a target. But it is a useful one, because it tells you two things immediately. Four hundred words is far more than most people produce when they are nervous and improvising. And it is nowhere near enough for four separate arguments, so trying to fit them all in is how answers end up as a list of unfinished thoughts.
What the 30 seconds is for
Not for drafting. You cannot write 400 words in 30 seconds and trying to will leave you reading a half sentence out loud. You can, however, write four or five words, and you are explicitly allowed to: handwritten notes on paper are permitted during the preparation time.
Use them for the skeleton, not the prose. Something like: issue, who it lands on, why, both sides, what I would do. Five words on a page is a map you can follow while the rest of your attention goes on speaking. A paragraph is something you will read out badly.
The scenario text stays on screen while you record, so you do not need to copy any of it down. That is 30 seconds you get back.
A structure that fits three minutes
Four movements, roughly 45 seconds each. The point is not that these proportions are correct in some deep sense, it is that you have a shape to be inside so you are never wondering what comes next while a camera runs.
- Name the thing. About 20 to 30 seconds. Say what the scenario is actually about in your own words, and say what makes it difficult. This is the part people skip because it feels obvious, and it is the part that tells a reviewer you understood the question rather than pattern matched it.
- Open it out. About 60 seconds. Who is affected and how differently. What is driving it, including the causes that sit above the individuals in the scenario. If there are two defensible positions, put both of them up properly, not as a strawman and a preference.
- Take a position. About 45 seconds. Say what you think and why. Say what would change your mind. UoA states you are not scored on your views, so the position is a vehicle for the reasoning, not the thing being marked, which is quite freeing once you believe it.
- Land it. About 20 to 30 seconds. One or two sentences that close the loop rather than trailing off. Watch the timer and start this before you have to.
For the Personal experience and Personal insight stations, the middle two movements change shape: you describe the specific thing that happened, then what it did to you and to others, then what you took from it. UoA asks you not to choose the worst thing that has ever happened to you, and that is practical advice rather than politeness. An anecdote you are still inside is very hard to be reflective about with a camera running.
The three things that go wrong on camera
Running out of things to say at ninety seconds
By far the most common. The fix is not more content in your head, it is a habit of going one layer down instead of moving on. You have named an effect: name who it falls on hardest. You have named a group: say why the system produces that result rather than a different one. Depth is what fills three minutes. Breadth runs out.
Realising the structure is wrong at forty seconds
You cannot restart, so do the thing you would do in conversation: say so and turn. UoA explicitly says that if you change your mind during the interview you should not be afraid to say so, and that being prepared to shift your position shows you are open to change and flexible in your thinking. That is permission, in writing, to think out loud.
Talking to nobody
There is no face to read, so people either speed up or flatten out. Practising in front of a mirror does not fix this because a mirror gives you feedback. Practising into a lens does, because a lens gives you exactly what the real thing gives you: nothing.
How to practise it so the practice counts
- Time everything, every time. An untimed rep teaches you the wrong pace, which is the specific skill this format tests.
- One take, always. The moment you allow yourself a second attempt you are practising a format that does not exist.
- Watch it back. Uncomfortable, and the only way to find out that you say "sort of" fourteen times or that you finished at 1:50.
- Do it on a laptop, not a phone. UoA says a tablet or mobile phone cannot be used on the day.
- Vary the domain. Seven stations means seven different kinds of question. Rehearsing one shape well is how people get caught by the Career choice station.
The seven domains and the eight attributes are here, and the full format, second by second, is here. If you want to do all of the above without paying anybody, we wrote that page too.
"The reviewers will recognise that you’ve only had 30 seconds to think about the scenario so they are not expecting you to cover everything, they are looking for your ability to think on your feet."
Which is the whole thing, really. You are not being asked for the complete answer. You are being asked for a real one, out loud, in three minutes, first time.
the only way to learn three minutes is to talk for three minutes
timed station, one take, watch it back, then a written read on your pacing and what you left out. free to start.
keep reading
- The Auckland MMI is asynchronous now: the Kira Talent format, explainedWhat actually happens at a station, second by second, and why most prep pages describe the wrong format.
- Auckland MMI 2026: the date, the day, and what to have readyOne sitting, one time, everybody at once. The dates and the device rules in one place.
- What the Auckland MMI actually scores you onSeven domains, eight attributes, and the rule that three are marked per station without telling you which.
- Auckland is dropping the UCAT: what changes for 2027 and 2028 entryGPA stops being a ranking lever, UCAT disappears, and Casper arrives at half the weight.
- Casper in New Zealand, and what it does to Auckland selectionhalf of the 2028 ranking, and almost nothing written for a New Zealand applicant.
read the official pages
everything on this page came from these. they are free, they are the only ones that count, and they change.
cutline is an independent study tool. we are not affiliated with, endorsed by, or connected to the University of Auckland, Kira Talent, or Acuity Insights. formats and requirements change, so always confirm the details against the official pages linked above before you rely on them.