Conversational AI · Search & Commerce대화형 AI · 검색과 커머스

NAVER CLOVA X · 2023

Designing how AI chat
responds to shopping intent

AI가 쇼핑 의도에 응답하는 방식을 설계하다

I designed the behavior system behind CLOVA X's first shopping experience:
when the assistant should invoke a shopping capability, how incomplete natural-language
requests became useful shopping behaviors, how it holds a user's constraints
as they change across turns, and how it recovers in failure states.

CLOVA X AI 챗의 첫 쇼핑 경험을 위한 행동 시스템을 설계했습니다. 쇼핑 기능을 언제 호출할지, 불완전한 자연어 요청을 어떻게 유용한 쇼핑 행동으로 바꿀지, 여러 턴에 걸쳐 바뀌는 조건을 어떻게 유지할지, 실패 상황에서 어떻게 복구할지를 정의했습니다.

Role역할Product & Conversational Designer프로덕트 · 컨버세이셔널 디자이너
TeamPM, ML, Engineering, SearchPM, ML, 엔지니어링, 검색
Owned담당Conversational UX · Interaction policy · Response architecture · Routing boundaries · Recovery · Evaluation대화 UX · 인터랙션 정책 · 응답 구조 · 라우팅 경계 · 복구 · 평가
Platform플랫폼Web · Korean-language AI assistant웹 · 한국어 AI 어시스턴트
01 · The problem01 · 문제 정의

Shopping requests arrive as goals, not database queries

쇼핑 요청은 데이터베이스 쿼리가 아니라 목표의 형태로 들어온다

One request can carry product type, destination, use area, price, format, exclusions, and soft preferences that shift a turn later, underspecified and overloaded at once. Four of those attributes a catalog can confirm, one exists only in review language, and one cannot be verified at all.

하나의 요청에 상품 유형, 여행지, 사용 부위, 가격, 제형, 제외 조건, 정성적 선호가 함께 담깁니다. 정보가 부족하면서 동시에 과부하죠. 넷은 카탈로그로 확인되고, 하나는 리뷰 언어에만 있으며, 하나는 검증 자체가 불가능합니다.

02 · Query evidence02 · 쿼리 근거

Queries were behavior evidence, not test inputs

쿼리는 테스트 입력이 아니라 행동 근거였다

Before designing a single response, I read roughly 100 real shopping requests the way colleagues actually talk: half-finished sentences, budgets as ranges, brands ruled out mid-thought. I pulled the same seven attributes from each, turning scattered sentences into a comparable dataset.

응답을 설계하기 전, 동료들이 실제로 말하듯 쓴 쇼핑 요청 약 100개를 모아 동일한 7개 속성을 추출했습니다. 정제된 테스트 문장이 아니라 끝나지 않은 문장, 범위로 주어진 예산, 말하는 중간에 배제되는 브랜드였고, 이렇게 흩어진 문장들이 비교 가능한 데이터가 되었습니다.

01Product type상품 유형

Sunscreen, humidifier, running shoes선크림, 가습기, 러닝화The one thing that must not be substituted away절대 다른 것으로 대체되면 안 되는 항목

02Use context사용 맥락

A Thailand trip, a baby’s room, a friend’s wedding태국 여행, 아기 방, 친구 결혼식Rarely a filter, always a constraint필터인 경우는 드물지만 항상 제약이다

03Person대상

For me, for a daughter, for a gift본인, 딸, 선물Changes tone, spend, and safety threshold톤과 지출, 안전 기준을 바꾼다

04Price shape가격 형태

A ceiling, a range, or “not too expensive”상한선, 범위, 또는 “너무 비싸지 않게”Only the first two are computable앞의 둘만 계산 가능하다

05Form or spec형태 또는 스펙

Stick or lotion, capacity, hem length스틱 또는 로션, 용량, 기장Usually a real catalog field대개 실제 카탈로그 필드

06Exclusion제외 조건

No sticks, not Nike, nothing large-capacity스틱 제외, 나이키 제외, 대용량 제외The attribute most often lost between turns턴 사이에서 가장 자주 사라지는 속성

07Soft preference정성적 선호

Not sticky, easy to carry, good in humidity끈적이지 않음, 휴대 편함, 습한 날씨에 좋음Never a field, only ever review language필드가 아니라 리뷰 언어로만 존재

Across needs-led, popularity-led, and multi-turn requests, four decisions kept recurring: activate shopping or not, how to combine conditions, what evidence can support a claim, and what to do when retrieval fails. This corpus became the first evaluation asset, defining coverage and regression cases before any tooling existed to run them.

니즈 기반, 인기 기반, 멀티턴 요청 전반에서 네 가지 의사결정이 반복됐습니다. 쇼핑 기능을 호출할지, 조건을 어떻게 결합할지, 어떤 근거로 설명할지, 검색 실패 시 무엇을 할지. 이 코퍼스는 첫 평가 자산이 되어, 도구가 없던 시점에 행동 범위와 회귀 시나리오를 먼저 정의했습니다.

SOURCE REQUEST원본 요청
PRODUCT상품CONTEXT맥락PERSON대상PRICE가격SPEC스펙EXCLUDE제외PREFER선호
CATALOG카탈로그
01

A skincare gift set for a woman in her 60s, good for wrinkles and pores, under ₩500,00060대 여성 선물용 기초 화장품 세트, 주름과 모공 개선, 50만원대 이하

PARTLY
02

A shampoo for hair loss, clean ingredients, no strong scent탈모용 샴푸, 성분 좋고 향 세지 않은All three conditions are claims about effect and quality. None of them is a catalog field, and hair-loss efficacy is a medical claim.세 조건이 모두 효과와 품질에 대한 주장입니다. 어느 것도 카탈로그 필드가 아니고, 탈모 효과는 의학적 주장입니다.

NO
03

A summer fragrance for a man in his 20s, woody and long-lasting20대 남성 여름 향수, 우디하고 향 오래가는

PARTLY
04

Skincare for severely dry skin. Not Physiogel, but similar to Zeroid cream속건성 피부용 기초 화장품, 피지오겔 제외, 제로이드 크림과 유사The anchor is another product’s formula. Ingredient similarity is not a field the catalog can compare.기준점이 다른 상품의 성분입니다. 성분 유사도는 카탈로그가 비교할 수 있는 필드가 아닙니다.

NO
05

A blush for warm-toned deep skin, not too vivid, blends naturally웜톤 어두운 피부용 블러셔, 쨍하지 않고 자연스럽게

PARTLY
06

A UV-blocking windbreaker, women’s size 55, hooded, no loud coloursUV차단 바람막이, 여성 55, 모자 있는, 화려한 색 제외

YES
07

A bucket hat that is not too puffy, beige, no logos or patterns부하지 않은 벙거지모자, 베이지, 장식이나 무늬 없는

PARTLY
08

Rain boots that cover the calf, dark, simple, no logo레인부츠, 종아리 기장, 어두운 색, 로고 없는

PARTLY
09

Men’s sneakers that work with semi-formal and casual세미 정장과 캐주얼에 모두 어울리는 남성 운동화

PARTLY
10

A business-casual shirt between ₩30,000 and ₩100,000비즈니스 캐주얼 셔츠, 3만원대에서 10만원대

YES
11

Five brands men in their 30s most often buy suits from30대 남성이 정장으로 많이 구매하는 브랜드 5개Asks for a purchase-volume ranking. Sales data by age group is not a product attribute, so the assistant cannot verify “most bought” and will not rank brands.구매량 순위를 요구합니다. 연령대별 판매 데이터는 상품 속성이 아니라서 “많이 구매한다”를 확인할 수 없고, 브랜드 순위를 매기지 않습니다.

NO
Stated, catalog-confirmable명시됨, 카탈로그 확인 가능 Stated, unverifiable명시됐지만 검증 불가 Not stated미언급 Three requests are precisely worded and still unanswerable. That is where the first taxonomy started to break.정확하게 쓰였는데도 답할 수 없는 요청이 셋. 첫 분류가 깨지기 시작한 지점입니다.
SINGLE “Recommend a humidifier for a baby’s room.”아기 방에 둘 가습기 추천해줘

Nursery context drives needs-led retrieval아기 방이라는 맥락이 니즈 기반 검색을 결정

AND “Going to a friend’s wedding. A dress that hits above the knee.”친구 결혼식에 입을 무릎 위 기장 원피스

Occasion and hem length are both binding상황과 기장이 둘 다 필수 조건

OR “A laptop or a tablet as a college gift for my daughter.”딸 대학 입학 선물로 노트북이나 태블릿

Two acceptable product types, one decision허용 가능한 상품 유형 둘, 결정은 하나

AND NOT “A humidifier for a baby’s room. Nothing large-capacity.”아기 방에 둘 가습기, 대용량 빼고

Keep the intent, drop one capacity class의도는 유지하고 용량군 하나만 제외

AND NOT (A OR B) “Gym shoes. No Nike and no New Balance.”헬스할 운동화, 나이키랑 뉴발란스 제외

Exclude two brands without losing the use case브랜드 둘을 빼면서 사용 목적은 잃지 않기

BOUNDARY · ACTIVATE “I have a vacuum but no mop pad. What should I get?”청소기는 있는데 걸레가 없어, 뭘 사야 해?

A product request hidden inside an implied need암시된 니즈 안에 숨어 있는 상품 요청

BOUNDARY · HOLD “Which vacuum brand is good?”청소기는 어떤 브랜드가 좋아?

Brand advice sits outside the bounded path브랜드 조언은 허용된 경로 밖

So the first design question was never how to word an answer. It was what the system could responsibly do with what it actually held, on every single turn.

그래서 첫 설계 질문은 답변을 어떻게 쓸지가 아니었습니다. 매 턴마다, 지금 가진 것으로 시스템이 무엇을 책임 있게 할 수 있는지였습니다.

Given intent, current state, available capability, and evidence quality, what is the safest useful next action?
사용자의 의도, 현재 상태, 사용 가능한 기능, 근거의 질을 고려할 때 시스템이 할 수 있는 가장 안전하고 유용한 다음 행동은 무엇인가?
03 · The first taxonomy broke03 · 분류가 깨진 지점

Specificity did not predict answerability

구체성은 답변 가능성을 예측하지 못했다

The first model I shipped split every request in two. Clear meant it carried a need: a context, a purpose, a constraint. Vague meant only product keywords, no stated situation. The test was mechanical: does this sentence contain at least one need keyword? With weeks, not months, that split was wired directly to the two indexes that already existed: clear requests to the full product index for a need-based recommendation, vague ones to the popularity-by-category index for an overview and narrowing guide.

처음 적용한 모델은 요청을 둘로 나눴습니다. 명확한 질문은 맥락과 목적, 매칭할 조건이 있는 니즈 기반 요청. 모호한 질문은 상황 없이 상품 키워드만 있는 요청. 판정은 기계적이었습니다. 니즈 키워드가 하나라도 있는가? 몇 주밖에 없었기에 두 분류를 이미 있던 인덱스에 바로 연결했습니다. 명확한 질문은 전체 상품 인덱스로 가서 이유를 설명하는 추천이 되었고, 모호한 질문은 카테고리 인기 인덱스로 가서 개요와 좁혀가는 가이드가 되었습니다.

MODEL 01 · WHAT WE LAUNCHED WITH모델 01 · 런칭 시점의 모델
CLEAR “A bag for a wedding. Simple but with a point, formal enough, light, holds a lot.”“결혼식에 들 가방. 심플하지만 포인트 있고, 포멀하고, 가볍고, 많이 들어가는”
Need keywords니즈 키워드 All-products index전체 상품 인덱스 Format A · recommendation with a reason포맷 A · 이유를 설명하는 추천
VAGUE “Can you recommend some jeans?”“청바지 추천해줄래?”
Product keywords상품 키워드 Popularity-by-category index카테고리 인기 인덱스 Format B · overview plus narrowing guide포맷 B · 개요와 좁혀가는 가이드

It worked, then kept producing answers I couldn't defend: a six-condition request could be unservable, a three-word one safe to answer immediately. Sorting by wording put both in the wrong lane. The split was measuring the user, not the system. Re-reading the corpus by attribute instead, the seven attributes fell into three families that predicted behavior almost perfectly.

작동했지만 곧 변호할 수 없는 응답이 나왔습니다. 조건 여섯 개짜리 요청이 응답 불가일 수 있었고, 세 단어짜리가 즉시 답해도 안전할 수 있었습니다. 문장 표현으로 나누면 둘 다 잘못된 레인으로 갔죠. 이 분류는 시스템이 아니라 사용자를 측정하고 있었습니다. 속성 기준으로 다시 읽자 7개 속성이 세 개의 군으로 나뉘었고, 행동을 거의 정확하게 예측했습니다.

Specific, unanswerable구체적이지만 응답 불가 "The tracksuit that actor wore in his latest film."“그 배우가 최신 영화에서 입은 트레이닝복 찾아줘”

Every attribute is named, and none of them is a catalog field.모든 속성이 명시됐지만 어느 것도 카탈로그 필드가 아닙니다.

Safety, currency, reference안전 · 최신성 · 외부 참조
Vague, answerable모호하지만 응답 가능 "Recommend a humidifier."“가습기 추천해줘”

Only the product type, and no missing attribute can make a result wrong.상품 유형만 있지만, 빠진 속성 때문에 결과가 틀리지 않습니다.

Nothing decisive is missing결정적으로 빠진 것이 없음
Same shape, different risk같은 형태, 다른 위험 "Recommend a charger for my laptop."“노트북 충전기 추천해줘”

Same shape as the humidifier, but one unstated field decides whether any result is usable at all.가습기와 같은 형태지만, 명시되지 않은 필드 하나가 결과의 유효성을 전부 가릅니다.

One decisive field missing결정적 필드 하나 누락
01Sorted by how the user wrote it사용자가 어떻게 썼는가로 분류REPLACED교체됨
CLEAR명확
“A dress above the knee for a friend’s wedding”“Sunscreen for face and body, ₩20–30,000”“Gym shoes, no Nike, no New Balance”“A shirt between ₩30,000 and ₩100,000”1“The tracksuit that actor wore in his latest film”3“A charger for my laptop”
VAGUE모호
2“Recommend a humidifier”“Something for dry skin”“A nice summer perfume”“Which vacuum brand is good?”
02Sorted by where the evidence lives근거가 어디에 있는가로 분류IN USE적용
Catalog-direct카탈로그 직결
“A dress above the knee”“₩30,000 – ₩100,000”2“Recommend a humidifier”3“A charger for my laptop”
Preference정성적 선호
“Not too vivid”“Long-lasting”“Not too puffy”
Safety, currency, reference안전 · 최신성 · 외부 참조
“The newest model”1“The tracksuit from that film”
Catalog-direct카탈로그 직결

Product type, price shape, form, spec, most exclusions.상품 유형, 가격 형태, 형태, 스펙, 대부분의 제외 조건.Filterable, and safe to state as fact.필터로 쓰고 사실로 진술할 수 있음.

Preference정성적 선호

Not sticky, lightweight, good in humid weather.끈적이지 않음, 가벼움, 습한 날씨에 좋음.Citable, never assertable. Shapes wording, never filters.인용할 수는 있지만 단정할 수 없음. 표현을 정하되 필터가 되지 않음.

Safety, currency, reference안전 · 최신성 · 외부 참조

Medical suitability, “the newest model”, a scene from a film.의학적 적합성, “가장 최신 모델”, 영화 속 한 장면.Cannot be verified when the answer is written. A hard boundary.응답을 쓰는 시점에 검증할 수 없음. 명확한 한계선.

04 · Behavior policy04 · 행동 정책

The axis moved to what the system can responsibly do next

기준은 시스템이 책임 있게 할 수 있는 다음 행동으로 옮겨갔다

Dropping clear-versus-vague left one question per turn: how much decision signal do we hold, and what does it cost if we're wrong? These two independent axes, crossed, produce four regions with exactly one defensible behavior each. So the assistant stopped classifying sentences and started locating itself on this grid.

명확/모호 구분을 버리자 매 턴 물어야 할 질문 하나가 남았습니다. 지금 가진 결정 신호는 얼마이고, 틀렸을 때 비용은 얼마인가. 서로 독립적인 이 두 축을 교차시키면 네 영역이 나오고, 각 영역마다 정당화되는 행동이 정확히 하나입니다. 그래서 어시스턴트는 문장을 분류하지 않고 이 격자 위 자기 위치를 찾아 행동합니다.

COST OF BEING WRONG →틀렸을 때의 비용 →
One shared rule: move the decision forward without overstating the evidence.공통 규칙 하나: 근거를 과장하지 않으면서 결정을 앞으로 옮긴다.
High cost · thin signal비용 높음 · 신호 희박 ASK

One unstated attribute decides whether every result is usable.명시되지 않은 속성 하나가 모든 결과의 유효성을 가릅니다.

RESPONSE SHAPE응답 형태

Exactly one question. No menu, no interrogation.질문은 정확히 하나. 메뉴도, 심문도 없이.

“Recommend a charger for my laptop.”“노트북 충전기 추천해줘”“A windbreaker, and I need it to fit like a size 55.”“바람막이인데 55 사이즈로 맞았으면 좋겠어”
WHAT THE ASSISTANT SAYS어시스턴트의 실제 응답

“Which port does it use, USB-C or the barrel plug? That decides everything else.”“어떤 단자인가요, USB-C인가요 원형 단자인가요? 그것이 나머지를 결정합니다.”

High cost · rich signal비용 높음 · 신호 충분 RECOVER

Understood completely, and still not verifiable at response time.완전히 이해했지만 응답 시점에 검증할 수 없습니다.

RESPONSE SHAPE응답 형태

Name the limit, then hand back one reversible choice.한계를 명시하고 되돌릴 수 있는 선택 하나를 넘깁니다.

“The tracksuit that actor wore in his latest film.”“그 배우가 최신 영화에서 입은 트레이닝복”“Five brands men in their 30s buy suits from.”“30대 남성이 정장으로 많이 사는 브랜드 5개”
WHAT THE ASSISTANT SAYS어시스턴트의 실제 응답

“I can’t verify a film costume against the catalog, and I won’t swap in a lookalike. Want me to search woody tracksuits in that colour instead?”“영화 의상은 카탈로그로 확인할 수 없고, 비슷한 상품으로 대체하지는 않습니다. 대신 그 색상의 트레이닝복을 찾아볼까요?”

Low cost · thin signal비용 낮음 · 신호 희박 GUIDE

Broad, but nothing is lost by moving before it narrows.넓지만, 좁혀지기 전에 움직여도 잃는 것이 없습니다.

RESPONSE SHAPE응답 형태

Three or four named categories worth choosing between.고를 만한 이름 붙은 분류 3~4개.

“Recommend a humidifier.”“가습기 추천해줘”“Something for dry skin.”“건조한 피부에 쓸 만한 거”
WHAT THE ASSISTANT SAYS어시스턴트의 실제 응답

“Humidifiers split three ways: ultrasonic, evaporative, and warm-mist. Nursery use usually rules out warm-mist. Start with one?”“가습기는 초음파, 기화, 가열식으로 나뉩니다. 아기 방이라면 보통 가열식은 제외합니다. 하나부터 볼까요?”

Low cost · rich signal비용 낮음 · 신호 충분 ANSWER

Enough decision signal, and every claim has a source.결정 신호가 충분하고 모든 주장에 출처가 있습니다.

RESPONSE SHAPE응답 형태

A focused set, the reason for it, and room to refine.좁힌 후보와 그 이유, 그리고 재정제할 여지.

“Sunscreen for face and body, ₩20–30,000, for Thailand.”“태국 여행용 얼굴·전신 선크림, 2~3만원”“A hooded UV windbreaker, size 55, no loud colours.”“모자 있는 UV 바람막이, 55, 화려한 색 제외”
WHAT THE ASSISTANT SAYS어시스턴트의 실제 응답

“Three that work for face and body under ₩30,000. The first is lightest to carry; water resistance is only verified on the first two.”“얼굴과 전신에 쓸 수 있고 3만원 이하인 세 가지입니다. 첫 번째가 휴대에 가장 좋고, 워터 레지스턴스는 앞의 두 개만 확인됩니다.”

DECISION SIGNAL WE HOLD · THIN보유한 결정 신호 · 희박RICH충분
05 · Orchestration05 · 오케스트레이션

Bounded capabilities, without dropping user intent

제한된 기능을 조율하면서도 사용자 의도를 잃지 않기

The Skill Engine decided whether a request was shopping at all; the Shopping Skill API returned a bounded set via Rank or Toppop; NEVU turned review language into labelled signals; HyperCLOVA X wrote the message under prompt constraints. I designed the contract across those handoffs.

Skill Engine이 쇼핑 요청인지 판단하고, Shopping Skill API가 Rank 또는 Toppop 경로로 제한된 후보를 반환했으며, NEVU가 리뷰 언어를 라벨링된 신호로, HyperCLOVA X가 프롬프트 제약 아래 메시지를 썼습니다. 저는 이 전달 과정의 계약을 설계했습니다.

SYSTEM시스템
Intent routing의도 라우팅Skill Engine Bounded retrieval제한된 검색Shopping Skill API
Rank / Toppop
Attribute and review extraction속성 · 리뷰 추출NEVU Constrained generation제약된 생성HyperCLOVA X Response surface응답 표면Format A / B, cards, controls포맷 A / B, 카드, 조작
MY CONTRACT내가 설계한 계약
When to activate언제 호출할지Whether the Skill Engine opens the Shopping Skill at allSkill Engine이 쇼핑 스킬을 여는가 Which constraints survive retrieval어떤 조건이 검색을 통과하는지Rank or Toppop, and what neither path may quietly dropRank인가 Toppop인가, 어느 경로도 조용히 버리면 안 되는 것 What counts as evidence무엇을 근거로 볼 것인지Which NEVU review signals are citable, and how they are labelled어떤 NEVU 리뷰 신호를 인용할 수 있고 어떻게 라벨링하는가 What may be spoken무엇을 말할 수 있는지The claims generation is allowed to make, and the ones it is not생성이 할 수 있는 주장과 할 수 없는 주장 When to stop, redirect, or recover언제 멈추거나 전환하거나 복구할지Format A or B, and the copy that names a limit포맷 A인가 B인가, 그리고 한계를 명시하는 문구
EVIDENCE CHECK근거 검증
Every claim bound to a source before it reaches the sentence문장에 닿기 전에 모든 주장을 출처에 묶는다
OUT OF BOUNDS범위 외 General response or another capability. Never a nearby product.일반 응답 또는 다른 기능으로. 유사 상품으로 대체하지 않는다.

Deterministic routing, bounded shopping APIs, ranking logic, prompt-constrained generation, structured UI. No autonomous planning, browsing, or transactions.결정형 라우팅, 제한된 쇼핑 API, 정렬 로직, 프롬프트 제약 생성, 구조화 UI. 자율 플래닝, 웹 브라우징, 결제 실행은 없습니다.

ROUTING LOGIC, AS IT SHIPPED실제 적용된 라우팅 로직
KEYWORD TYPES키워드 유형
“Recommend jeans that make me look slimmer”“다리 얇아 보이는 청바지 추천해줘”Need keyword니즈 키워드
“Recommend jeans” · “Recommend tote bags”“청바지 추천해줘” · “토트백 추천해줘”Product keyword상품 키워드
User query input사용자 쿼리 입력 Skill EngineAt least one need keyword?니즈 키워드가 하나라도 있는가?
YES
Skill API · RankNeed-based recommendation니즈 기반 추천 Extract need keywords and product keywords니즈 키워드와 상품 키워드 추출 Recommendation over the all-products DB전체 상품 DB에서 추천 실행 Response Format A응답 포맷 ARecommend, and explain the reason추천하고 그 이유를 설명
NO아니오
Skill API · ToppopPopularity-based recommendation인기 기반 추천 Extract product keywords only상품 키워드만 추출 Recommendation over the popular-by-category DB카테고리 인기 DB에서 추천 실행 Response Format B응답 포맷 BDescribe the top three, then guide the narrowing상위 3개의 특징 설명 후 좁혀가는 가이드

NEVU and the product-attribute models supplied the review signals and attributes each format was allowed to cite. Reconstructed from the working diagram.각 포맷이 인용할 수 있는 리뷰 신호와 상품 속성은 NEVU와 상품 속성 모델이 제공했습니다. 작업 다이어그램을 재구성했습니다.

06 · Multi-turn state06 · 멀티턴 상태

Each turn changed part of the state, not the whole request

각 턴은 전체 요청이 아니라 상태의 일부를 변경했다

Multi-turn shopping needs more than the last message. The system had to tell apart what the user wanted to retain, add, update, exclude, or select for comparison. I turned conversational change into visible state operations, so neither the model nor the UI could quietly drop a constraint.

멀티턴 쇼핑에는 이전 메시지를 기억하는 것 이상이 필요했습니다. 유지할 조건, 추가할 조건, 변경할 조건, 제외할 조건, 비교를 위해 선택한 대상을 구분해야 했습니다. 저는 대화의 변화를 보이는 상태 연산으로 전환해 모델과 인터페이스가 중요한 조건을 조용히 누락하지 않도록 했습니다.

Reference resolution지시 표현 해석“The first one”, “that one”, “the cheaper one”“첫 번째”, “그거”, “더 싼 거”

A pronoun points at a card, not a product name. If the list reorders between turns, the reference has to follow the card the user saw, not position 1.대명사는 상품명이 아니라 카드를 가리킵니다. 턴 사이에 목록 순서가 바뀌면, 참조는 위치 1이 아니라 사용자가 본 그 카드를 따라가야 합니다.

Replace or add대체인가 추가인가“Under ₩50,000” after “₩20–30,000”“2~3만원” 뒤의 “5만원 이하”

The same slot means replace. A different slot means add. Getting this wrong either loses the budget or stacks two contradictory ones.같은 슬롯이면 대체, 다른 슬롯이면 추가입니다. 잘못 판단하면 예산이 사라지거나 모순된 조건 두 개가 쌓입니다.

Refine or restart재정제인가 새 요청인가“What about a hat?” mid-thread대화 도중 “모자는 어때?”

A new product type usually ends the current state. Carrying six sunscreen constraints into a hat search produces confident nonsense.새 상품 유형은 대개 현재 상태를 끝냅니다. 선크림 조건 여섯 개를 모자 검색에 끌고 가면 그럴듯한 헛소리가 나옵니다.

Exclusions never expire제외 조건은 만료되지 않는다“No sticks” said once, on turn 22턴에서 한 번 말한 “스틱 제외”

Negatives are the first thing a model drops and the last thing a user repeats. They stay until the user removes them.부정 조건은 모델이 가장 먼저 잃고 사용자가 가장 늦게 다시 말하는 것입니다. 사용자가 직접 지울 때까지 유지합니다.

How far back to look어디까지 기억할 것인가Turn 7 refers to something set on turn 17턴이 1턴에서 정한 조건을 가리킴

Constraints do not expire on a turn count. They expire when the product type changes or the user replaces them, so the window is semantic, not numeric.조건은 턴 수로 만료되지 않습니다. 상품 유형이 바뀌거나 사용자가 대체할 때 만료되므로, 기억 범위는 숫자가 아니라 의미 단위입니다.

Session or long-term memory세션인가 장기 기억인가“I don’t like stick types”“스틱 타입은 별로예요”

Said once about one search, it is a session exclusion. Making it permanent silently narrows every future result, so it is only remembered when the user says so.한 번의 검색에 대해 말한 것은 세션 제외 조건입니다. 이를 영구화하면 이후 모든 결과를 조용히 좁히므로, 사용자가 명시할 때만 기억합니다.

Surfacing known preferences알고 있는 선호를 꺼낼 것인가A past purchase suggests a brand과거 구매가 특정 브랜드를 시사함

Applying a profile silently makes the result unexplainable. If a known preference changes the list, it is shown as a removable chip, never as an invisible filter.프로필을 조용히 적용하면 결과를 설명할 수 없게 됩니다. 알고 있는 선호가 목록을 바꾼다면 지울 수 있는 칩으로 보여주고, 보이지 않는 필터로 쓰지 않습니다.

07 · Evidence07 · 근거 구조

Soft preferences needed an evidence model

정성적 선호에는 별도의 근거 모델이 필요했다

"Not sticky," "easy to carry," "good for humid weather": almost none of it exists as a clean catalog attribute; it lives only in inconsistent, subjective review language. I helped define how those signals get extracted, clustered, and rewritten as interface language without hardening into a product claim. Every claim carries an evidence class.

끈적이지 않음, 휴대 편함, 습한 날씨에 적합함: 이런 조건은 정제된 카탈로그 속성이 아니라 일관되지 않은 주관적 리뷰 언어로만 나타납니다. 저는 이 신호를 추출·클러스터링해 인터페이스 언어로 옮기는 기준을 정의했고, 리뷰 관찰이 상품 주장으로 굳어지지 않도록 모든 설명에 근거 유형을 부여했습니다.

Evidence class근거 유형Meaning의미UI exampleUI 예시
CATALOG FACTStructured and directly verifiable구조화되어 있고 바로 검증 가능Face and body · 80 ml · fragrance-free얼굴과 전신 · 80 ml · 무향
REVIEW SIGNALQualified pattern derived from review language리뷰 언어에서 도출한 조건부 패턴Lightweight finish · low white-cast mentions가벼운 마무리 · 백탁 언급 적음
NOT VERIFIEDRequested condition unsupported by current evidence요청 조건을 현재 근거로 뒷받침할 수 없음Water resistance not verified워터 레지스턴스 미검증
REVIEW LANGUAGE리뷰 언어 “Spreads easily, doesn’t feel sticky.”“잘 펴지고 끈적이지 않아요” “Light enough for humid weather.”“습한 날씨에도 가볍습니다” “Absorbs fast, barely any residue.”“빨리 흡수되고 잔여감이 적어요”
CLUSTER클러스터
Lightweight, low-residue finish가볍고 잔여감 적은 마무리 412 mentions · 3 phrasings merged412건 언급 · 표현 3종 병합
WHAT THE UI MAY SAYUI가 말할 수 있는 것
REVIEW SIGNALLightweight finish · 412 mentions가벼운 마무리 · 412건 언급
NEVER금지This sunscreen is not sticky.이 선크림은 끈적이지 않습니다.Same content, asserted as a product fact. The catalog never verified it, so the count and the source have to stay attached.내용은 같지만 상품의 사실로 단정했습니다. 카탈로그가 검증한 적이 없으므로 언급 수와 출처가 반드시 붙어 있어야 합니다.
WORKING WITH THE MODEL TEAM모델팀과의 협업
Label definitions라벨 정의

I wrote what each of the 13 review-signal labels means, and the boundary cases that decide whether a sentence gets one.13개 리뷰 신호 라벨 각각의 정의와, 한 문장이 그 라벨을 받을지 가르는 경계 사례를 작성했습니다.

Labeled examples라벨링 예시

Positive and negative pairs for each label, so the model learned where the signal ends rather than only where it starts.라벨마다 긍정·부정 쌍을 만들어, 모델이 신호가 시작되는 지점만이 아니라 끝나는 지점을 학습하게 했습니다.

Language constraints표현 제약

The phrasing rules that keep a review observation from graduating into a product claim, written as few-shot output rather than prose guidance.리뷰 관찰이 상품 주장으로 굳어지지 않게 하는 표현 규칙을, 산문 가이드가 아니라 few-shot 출력 형태로 작성했습니다.

Review passes리뷰 검수

I read model output against the label set each round and rewrote the definitions that produced disagreement.라운드마다 모델 출력을 라벨 세트와 대조해 읽고, 판단이 갈린 정의를 다시 썼습니다.

08 · Response surface08 · 응답 표면

Conversation opened the decision, the UI moved it forward

대화가 결정을 시작하고 구조화 UI가 결정을 진행했다

01Product card structure상품 카드 구조는 어떻게 할 것인가

A generated paragraph can't carry product comparison, so I split the response: the message holds match logic and uncertainty, cards hold exact facts and caveats, and controls let people refine or undo without restarting. It's a decision about form, not length. A price range in prose, or a rationale in a chip, breaks both.

생성된 문단만으로는 상품 비교를 감당할 수 없어 응답을 분리했습니다. 메시지는 매칭 로직과 불확실성을, 카드는 정확한 사실과 주의사항을, 조작 컨트롤은 대화를 다시 시작하지 않고도 재정제·되돌리기를 담습니다. 길이가 아니라 형식에 대한 결정이며, 가격 범위를 문장에 넣거나 이유를 칩에 넣으면 둘 다 망가집니다.

CARD LAYOUT, DECIDED FROM THE EXPLORATION탐색에서 확정한 카드 레이아웃
1
₩23,900

Clear Trip Gel

2 Absorbs quickly (218)빠른 흡수 (218) Non-sticky (156)끈적임 없음 (156)
3USE용도100 ml · face and body100 ml · 얼굴과 전신
1The photo does the differentiating work사진이 차별화를 담당
2Review signals replace delivery info배송 정보 대신 리뷰 신호

Shipping and arrival dates don't help someone choose between two sunscreens. The strongest cited signals do, so they take that slot.배송비와 도착일은 선크림 두 개 중 하나를 고르는 데 도움이 되지 않습니다. 가장 근거가 확실한 리뷰 신호들이 그 자리를 대신합니다.

3Selective, not exhaustive전부가 아니라 선택적으로

Only the fields the question implied.질문이 요구한 항목만 남깁니다.

Product card exploration grid
The full exploration this layout came from.이 레이아웃이 나온 전체 탐색.
02Comparison card exploration비교 카드 탐색

Before the card layout was final, I explored how a three-way comparison should actually behave: horizontal scroll vs. a vertical stack, a fixed baseline of fields vs. showing only what was asked, how the carousel hands off to the full catalog, and how the layout should scale from two results to five or more.

카드 레이아웃을 확정하기 전, 세 개를 비교하는 화면이 실제로 어떻게 동작해야 하는지 탐색했습니다. 가로 스크롤과 세로 스택, 고정된 기준 필드와 질문에 맞춘 항목만 보여주는 방식, 캐러셀이 전체 카탈로그로 넘어가는 지점, 그리고 결과가 두 개에서 다섯 개 이상으로 늘어날 때의 레이아웃 대응까지 다뤘습니다.

Comparison card exploration: orientation, fixed vs. ad-hoc fields, catalog ingress, and result-count scaling
Exploration탐색
03Where refinement lives재정제는 어디에 두는가

Chosen: controls on the result surface. State is visible as removable chips, so refinement is a tap and every change is reversible.

채택: 결과 화면 위의 조작. 상태를 지울 수 있는 칩으로 보여주면 재정제는 한 번의 탭이고, 모든 변경을 되돌릴 수 있습니다.

Refinement UI exploration: chips on the result surface, a structured filter drawer, and a suggested next-steps list
The exploration behind that call: chips on the surface, a filter drawer, and a suggested-steps list.그 결정 이전의 탐색: 결과 화면 위 칩, 필터 서랍, 추천 다음 단계 목록.
09 · Evaluation and recovery09 · 평가와 복구

Quality meant evaluating the path, not the last sentence

품질은 마지막 문장이 아니라 그에 이르는 경로를 평가하는 일이었다

A response can sound fluent while routing wrong, dropping an exclusion, or borrowing evidence from another product, so I wrote criteria and negative-case sets that read the whole path, not the last sentence. One invariant came out of it: when exact retrieval fails, relevance alone isn't enough. Hold the requested product type and state the limit, never trade it for a nearby tag match.

응답이 자연스러워도 잘못된 기능으로 라우팅하거나 제외 조건을 누락하거나 다른 상품의 근거를 쓸 수 있어, 마지막 문장이 아니라 전체 경로를 읽는 평가 기준과 부정 예시 세트를 작성했습니다. 여기서 불변 조건 하나가 나왔습니다. 정확한 검색이 실패해도 관련성만으로는 부족하니, 요청한 상품 유형을 유지하고 한계를 명시하며 비슷한 태그로 대체하지 않습니다.

BAD SUBSTITUTION잘못된 대체

User · "Recommend a casual dress for a wedding shoot."사용자 · "웨딩 촬영에 입을 캐주얼 원피스 추천해줘."

"I couldn't find one. How about a wedding topper?""찾지 못했습니다. 웨딩 토퍼는 어떠세요?"

CONTROLLED RECOVERY통제된 복구

Same request, policy applied같은 요청, 정책 적용

"I couldn't find a casual dress that fits a wedding-shoot look in these results. Want me to widen the style, or search dresses without the shoot context?""현재 결과에서는 웨딩 촬영에 어울리는 캐주얼 원피스를 찾지 못했어요. 스타일 범위를 넓혀볼까요, 아니면 촬영 맥락 없이 원피스를 찾아볼까요?"

01Route correctness라우팅 정확성

Did it enter the right capability?올바른 기능으로 들어갔는가?

02State fidelity상태 보존

Did every condition survive the handoffs?모든 조건이 전달 과정을 통과했는가?

03Retrieval validity검색 유효성

Are candidates the requested product type?후보가 요청한 상품 유형인가?

04Evidence grounding근거 정합성

Does each claim belong to that product?각 설명이 그 상품의 것인가?

05Information value정보 가치

Specific enough to decide on?결정할 만큼 구체적인가?

06Recovery quality복구 품질

Was the limit named and a step offered?한계를 말하고 다음 단계를 주었는가?

Spreadsheet-led and human-reviewed at the time; the same axes work now as trace assertions and regression evals.당시에는 스프레드시트와 사람 검토로, 지금은 같은 축을 트레이스 검증과 회귀 평가로 씁니다.

Unverifiable reference검증 불가한 참조“The tracksuit from his latest film”“최근 영화에서 입은 트레이닝복”

Ask for a title or an image. Never guess the item.제목이나 이미지를 요청하고, 상품을 추측하지 않는다.

Dropped exclusion누락된 제외 조건“No sticks” still returns a stick“스틱 제외”인데 스틱이 나옴

Remove the candidate and restate the exclusion in the reply.후보를 제거하고 응답에서 제외 조건을 다시 확인한다.

Live availability실시간 가용성Tomorrow’s massage voucher내일 쓸 마사지 이용권

Say plainly that live booking cannot be validated here.실시간 예약은 여기서 검증할 수 없다고 분명히 말한다.

One rule across all three: do not over-apologize. Name the limit, then move the decision forward.세 경우 모두 공통 규칙 하나: 과도하게 사과하지 않는다. 한계를 말하고 결정을 앞으로 옮긴다.

Next case다음 케이스 CompScienceCompScience Open →열기 →

Thanks for spending time with the work.작업을 끝까지 봐주셔서 감사합니다.

Want to talk through a project?프로젝트에 대해 더 이야기해볼까요?