Interaction design · Speculative research

UC Berkeley Master's thesis · 2024–2025

Designing an AI-led experience across voice, touch, and body

What makes us human? Future CAPTCHA began as an authentication problem and became a research-led interactive test about the qualities people use to describe being human. I designed the research, interaction choreography, voice behavior, interface, and working iOS prototype, then observed the experience with public visitors at exhibition scale.

RoleSolo researcher · designer · developer
ResearchDiary · survey · interviews / media
MediumInteractive iOS experience
RecognitionInternational Design Award GoldGlobal AI Design Award GoldUC Berkeley thesis honorable mentionFeatured at SF Design Week 2025
00 · Try it

Before I explain it, experience it.

About three minutes on the exhibition floor: the evaluator speaks, a visitor answers out loud, moves, and receives a score. Sound on, the voice is half the interface.

Recorded at the thesis exhibition, UC Berkeley, December 2024.The tester's voice and their taps happen off-frame, so what you are watching is the system's side of the conversation: the prompts, the listening state, the transcript coming back, and the score it settles on.

01 · Research question

I started with authentication. I ended up questioning the thing we were trying to authenticate.

My thesis began with telephone fraud. I spent months on it, mapping who gets targeted, why shame keeps victims silent, and where an intervention could sit. Then my attention shifted. The version of the problem that kept pulling at me was no longer a stranger on the phone but a synthesized voice that sounded like someone you love.

Where the question came from

Topic exploration board on voice scams: vulnerable groups, AI automated calls and deepfakes, psychological effects
Problem-space exploration: who is most vulnerable, how AI-generated calls change the scale, and why the psychological damage outlasts the financial loss.
Media research: a deepfake comparison video and a news segment on AI-imitated voices
Media research. Deepfake heists and AI-cloned voices moved the public conversation from prevention to detection.
Headlines: ChatGPT broke the Turing test, and a report on identifying AI-generated political voices
Desk research. Detection researchers, including Hany Farid at Berkeley, point to imperfections in real speech, and cloning keeps getting better at reproducing them.

Each detector eventually loses, which left me with a question underneath the whole effort.

What exactly are we trying to prove when we prove that someone is human?

That question moved the thesis from a real-world problem to a philosophical one: voice scam, to authentication, to humanness. Seven years of design education and practice had trained me to look for the practical solution first. For a master's thesis I wanted the opposite, the bigger and slower question that shapes how designers work with emerging technology.

Thesis journey map: voice scam to authentication to humanness, with the questions raised at each stage
Thesis journeyVoice scam, to authentication, to humanness, and the question each stage opened.
A 2x2 map of six concept directions, plotted from currently applicable to conceptual and from reactive to preventive
Choosing a directionA one-pager for each of six ideas, placed on a 2x2 matrix so I could prioritize the core problem.
02 · Field research

What makes us human? Whose definition of human?

A quick interview would have forced people to answer an abstract question on demand. I wanted three kinds of evidence: what people notice in their own lives, what a broader community intuitively associates with humanity, and how different disciplines frame the human-machine boundary.

Participants documented moments when they felt deeply human, with context, action, and feeling. The mundane answers became the most revealing: morning coffee, guacamole, being overwhelmed, small rituals, procrastination, relationships.

Lived experience

14-person diary study

Three days to a week, so the thought could develop inside ordinary life instead of on the spot.

Collective intuition

Campus-wide survey

I moved the question rather than assuming one community represented Berkeley: Soda Hall, Sutardja Dai, Life Sciences, Moffitt.

Expert & cultural lenses

Interviews + literature

Social science researchers, ML engineers, and a review of how media frames the boundary.

The survey board filling with responses, frame 1 of 7 The survey board filling with responses, frame 2 of 7 The survey board filling with responses, frame 3 of 7 The survey board filling with responses, frame 4 of 7 The survey board filling with responses, frame 5 of 7 The survey board filling with responses, frame 6 of 7 The survey board filling with responses, frame 7 of 7 01 / 07
Campus-wide surveyThe same board, carried between buildings and refilled as it went. Responses accumulated from a handful of notes to a full sheet.
Diary study materials
Diary studyReturned booklets from fourteen participants, three days to a week each.
03 · From findings to test

Research findings became interaction requirements.

I resisted forcing the research into a neat definition too early. Across diaries, interviews, survey responses, and media research, recurring clusters formed around five qualities, and each one had to become something a person actually does in the room.

A screen full of survey questions would have contradicted all of it. If embodiment mattered, participants needed to move. If emotional complexity mattered, fixed-choice responses were not enough. If unpredictability mattered, the system needed to observe behavior rather than ask people whether they were unpredictable.

What makes us human?

Handwritten diary entry: Love and nostalgia, mixed with sadness
Love and nostalgia, mixed with sadness
Handwritten diary entry: Contradictions, and trying to be at peace with them
Contradictions, and trying to be at peace with them
Handwritten diary entry: Having emotions, and reflecting on them
Having emotions, and reflecting on them
Survey board note: Compassion, empathy, forgiveness, smiles
Compassion, empathy, forgiveness, smiles
Handwritten diary entry: Raking leaves, using my muscles
Raking leaves, using my muscles
Handwritten diary entry: The cold, and a burned tongue
The cold, and a burned tongue
Handwritten diary entry: An ache in the legs after working out
An ache in the legs after working out
Handwritten diary entry: My body tells me what I need
My body tells me what I need
Handwritten diary entry: Not rational, not logical, no recipe
Not rational, not logical, no recipe
Handwritten diary entry: Thinking that language cannot explain
Thinking that language cannot explain
Survey board note: Subjectivity, and hope as opposed to pure logic
Subjectivity, and hope as opposed to pure logic
Survey board note: Laughter, and answers nobody expected
Laughter, and answers nobody expected
Handwritten diary entry: A slice of apple pie
A slice of apple pie
Survey board note: Indecision, inconsistency, and not being sure at all
Indecision, inconsistency, and not being sure at all
Survey board note: Making a mistake and getting back up
Making a mistake and getting back up
Handwritten diary entry: A sense of community, and of nature
A sense of community, and of nature
Handwritten diary entry: A soul, unique to each person
A soul, unique to each person
Handwritten diary entry: Something AI cannot create, only reorganize
Something AI cannot create, only reorganize
Survey board note: Gratitude for all the things in life
Gratitude for all the things in life
The project idea board: research question, desk research, disciplines, and the first synthesis clusters
Idea boardKept on the wall and updated as evidence arrived.
Survey responses clustered into labelled groups of sticky notes
Survey synthesisEvery note from the campus board, sorted into named clusters: empathy, memory, imperfection, belief, the unknown.
Digital synthesis board combining diary study, literature review, and interview data
Qualitative synthesisDiary study, literature and media review, and in-depth interviews merged into one board so the themes could be compared across sources.
04 · Experience choreography

Multimodal design became attention design.

Before building the prototype, I physically mapped the sequence of prompts, inputs, measurements, feedback, and final reflection. I was designing the rhythm of attention, not just individual screens: listen, respond, move, receive feedback, continue.

Not every human response should use the same input. The screen is the stimulus and touch is the answer when reaction should happen before explanation. Voice carries expression when the response is a value judgment. The body becomes the input when the quality being tested is embodiment, and once movement began, voice mattered more than text, because looking back at instructions would break the interaction.

Screen-first · touch

Visual affinity

Screen is the stimulus, touch is the answer, voice is the acknowledgment.

Voice-first · speech

Value response

Screen holds context; the visual layer only reports that the system is listening.

Body-first · movement

Embodiment

Continuous guidance moves into voice so nobody has to look back at the screen.

Physical experience map prototype
Physical experience mapThe interaction sequence, folded and rearranged before a line of code.
Presenting the physical experience map at a studio critique
Studio critiqueWalking the committee through the sequence before building the prototype.
05 · Voice & turn-taking

The evaluator needed authority without becoming inhuman.

I moved away from the bright, friendly, often feminine assistant archetype. The premise was closer to an airport checkpoint: the system was evaluating you. I worked through Google Cloud's voice catalog, listening to hundreds of them, before choosing a lower-pitched male voice that balanced authority, warmth, dry humor, and audibility in a noisy room. Casting was only half of it: the same voice at default settings still sounded like an assistant, so I slowed the delivery and dropped the pitch until it read as an evaluator.

01 · Casting the voice

Rejected · bright assistant Helpful, and completely wrong
{{ waveA }}

Warm and eager. It made the test feel like a product feature instead of an evaluation, and it disappeared under room noise.

Chosen · low, dry, deliberate Authority with a sense of humor
{{ waveB }}

Lower pitch, slower rate, flat delivery. Closer to an airport checkpoint than an assistant, and it still carried across a crowded room.

Waveforms are illustrative. Attach the two exported clips to turn this into an A/B listen.

The synthesis request: voice en-US-Casual-K, speaking rate 0.9, pitch minus 3
The settings that finished the castingSpeaking rate down to 0.9 and pitch down three semitones. Small numbers, and the difference between an assistant and an evaluator.

02 · Writing the line so it can be heard

WRITTEN

You have a name as an individual. Are you a real human? This is a mandatory test identifying humans among machines.

SPOKEN

Okay okay, [600ms] well...... You have a name as an individual. [700ms] So.. Jade, are you uh... real human? [800ms] This is a mandatory test identifying humans among machines, or, something in the middle.

Pauses and fillers live in the script itself, so every visitor hears the same performance.

03 · Who is holding the turn

VOICE
TEXT
MIC

Voice and text finish together. Listening opens after a deliberate gap, never on the final syllable.

Before

Silence after the prompt

Nothing on screen while the microphone was live. People asked out loud whether it was listening, then answered too quietly to register.

After

The transcript is the feedback

The recognizer streams its transcript and the screen renders it as it arrives. People corrected themselves without being told to: they saw a garbled line, stepped closer, said it again.

Micro-personality: the same tap, four different replies

A liked photograph gets "Nice..", a liked AI image gets "Interesting choice..", a rejected AI image gets "Okay," and a rejected photograph gets "dislike... noted." Neutral returns nothing at all.

Their job is character, not information. That asymmetry made the system feel as if it were watching rather than recording, and a flat delivery kept it from steering the next answer.

06 · Building it

I wanted to read emotion from the voice itself. That failed.

The original plan sent raw voice to Hume AI, a vocal emotion model. In exhibition conditions it was unreliable, so the working system transcribes speech on device and interprets the transcript semantically instead.

This was not an equivalent replacement. Vocal emotion asks how the person sounded; semantic emotion asks what the person expressed. I kept that limitation visible in the framing rather than pretending the system could measure something it no longer could. I changed the signal, not the interaction intent.

Speech in

SFSpeechRecognizer / AVFoundation

Interpretation

GPT-4, context and semantic emotion

System voice

Google Cloud TTS, scripted with SSML pauses

Movement

Vision Framework / ARKit

Emotion result screen
Result as interpretationBecause the fallback analyzed transcribed content rather than vocal affect, the output had to read as a reading, not a diagnosis.
07 · Exhibition

The final interface had to work without me explaining it.

This was not a lab prototype. An academic committee and hundreds of visitors encountered it while other people were talking, waiting, watching, and moving around them. No formal onboarding, ambient noise, a social cost to speaking in public, shifting attention, and spectators who made the interaction partly performative.

So the exhibition became a usability study in the wild. I stood to the side and watched, and these are the things I could not have learned in a quiet room.

Eight things I watched happen

  1. 01Spoke over the prompt, or waited too longnothing marked when listening began
  2. 02Missed what the evaluator saidaudio alone does not survive a loud room
  3. 03Asked a neighbor to repeat itlong prompts need a second pass
  4. 04Assumed it was brokensilence during processing read as failure
  5. 05Mumbled with an audience behind themquiet answers under-registered
  6. 06Turned back to read instructionsreading stopped the movement being measured
  7. 07Waited for someone else to go firstno proof the tracking was live
  8. 08Compared results out loudthe argument was the point, so it stayed

Believable enough to engage. Questionable enough to discuss.

The goal was never an accurate measurement of humanity. I considered the experience successful when participants stopped treating the machine as an authority and started debating the assumptions behind its judgment: why did it think that, could I fool it, would another person get the same result, and why do we need to prove that we are human at all.

Visitor taking the Future CAPTCHA test at the exhibition
A visitor taking the test while others wait and watch.
Future CAPTCHA result screen
The result screen.

Recognition

GoldInternational Design Award2025 GoldGlobal AI Design Award2025 Honorable mentionUC Berkeley thesis2024 FeaturedSF Design Week2025
08 · Take it yourself

A web version, rebuilt so you can try it here.

This is where those observations landed. The exhibited build used a tablet, a cast voice, capacitive touch, and head tracking, and the browser has none of that, so the port trades fidelity for access: the emotion reading is coarser, a pulse stands in for capacitive sensing, a cursor for head tracking, and your device speaks with whatever voice it has.

What it does carry is the script, the images, the scoring, and every interface change the exhibition argued for. Sound on. Five steps, about three minutes.

Everything spoken is also written

The evaluator's line appears on screen as it is said, so a missed word is never a lost instruction.

Listening is a state, not a guess

An explicit prompting, listening, captured, processing sequence, and the microphone opens after the prompt rather than on its last syllable.

Answers are shown back

The transcript streams while you speak, so you can see what was heard and say it again if it came out wrong.

Short, sequential instructions

One thing to do at a time instead of a paragraph to retain.

The score still refuses to explain itself

The one thing I did not fix.

Still no way to replay a prompt

Deliberate. The premise is an airport checkpoint, a test administered by someone standing in front of you, and you do not get to ask an officer to say it again from the top. A replay button would have made it a media player instead.

Visitors gathered around the installation at the exhibition, laughing as someone takes the test
A visitor watching her own movement tracked on the screen
Two visitors reading an emotion result on the screen
A crowded room at the thesis exhibition

The test was never going to measure humanity accurately.It was built to open up the discussion, and the lineeach person drew between human and machinerevealed their own definition, not the system's accuracy.

Next case CLOVA X 01 →

Thanks for spending time with the work.

Want to talk through a project?