Skip to main content

Agent Experiences ki Design

Woh surface jahan insan khud kaam karne wali machine par bharosa seekhta hai

Aap ne is kitab mein aisi machines banai hain jo kaam karti hain: general agents, Digital FTEs aur aisi workforces jo apne colleagues khud hire karti hain. Yeh course us patli lekin faisla-kun layer ke bare mein hai jahan insan machine se milta aur faisla karta hai ke is par bharosa karna hai ya nahin.

Woh layer ab screen nahin hai. Jab software sirf commands ka jawab deta tha, design ka matlab buttons ko is tarah rakhna tha ke insan machine chala sake. Jab software khud action leta hai, insan driver nahin rehta. Woh kaam delegate karta hai. Aur delegation transaction nahin, relationship hai.

Is liye discipline ka naam badalta hai. Hum insan ke chalaye hue interface ki design se insan ki supervise ki hui partnership ki design tak jate hain. Craft ab "button aasani se mil jaye" nahin, balkeh "machine ka judgment samajh aaye, autonomy adjust ho sake aur ghalti se recovery mumkin ho" hai. Yeh course isi cheez ki taleem deta hai.

Aik jumle ka markazi khayal

Agentic product ke aik waqt mein do users hote hain: aik insan jise is par bharosa karna hai aur doosre agents jinhein ise parse karna hai. Aap ka kaam aisa surface design karna hai jo kisi aik ko dhoka diye baghair dono ki zaroorat poori kare.

Aap kya banayeinge. Aakhir tak aap kitab mein pehle banaye gaye aik Digital FTE ka Agent Experience Brief draft kar chuke honge: human-facing trust surface, agent-facing machine surface, autonomy ladder aur recovery plan. Naya code nahin; Human-Agent Teams mein operating documents ki tarah aap agent ko brief banane ki hidayat deinge. Appendix C mein blank fillable version hai. Hands-On Lab mein aap apna pehla MCP App bhi ship kareinge: refund Worker ka working approval widget, jo official create-mcp-app skill rakhne wale coding agent, Claude Code ya OpenCode, ko direct karke banega. Leaders aur designers ke liye concepts ke darmiyan Reader track bhi hai jismein build karna zaroori nahin.

Chaar parts mein atharah concepts hain: tabdeeli, human surface, machine surface aur nai craft. Parhne mein taqreeban do ghante; capstone brief aik shaam ka kaam aur Reader track focused aik ghanta hai. Worked Example ke baad Hands-On Lab Part 3 ko working code banata hai: coding agent ko direct karke apna pehla MCP App banayein aur ship karein. Aakhir ke appendices MCP Apps ki anatomy aur MCP Apps muqable OpenAI Apps SDK ka sawal map karte hain.

Alfaaz naye hain? Saada zaban mein das lafz

Yeh course chand technical alfaaz baar baar istemaal karta hai. Aage koi lafz na roke, is liye unhein aik baar saada zaban mein samjhein:

  • Agent: aisa software jo sirf sawal ka jawab dene ke bajaye diye gaye goal ke liye khud actions leta hai.
  • Worker / Digital FTE: haqeeqi job ke liye banaye agent ka is kitab mein naam, jaise software se bana customer-support employee.
  • MCP (Model Context Protocol): open standard jo agents ko tools discover aur call karne deta hai; agents aur software ke darmiyan universal plug.
  • Connector / MCP server: woh plug, yani software ka hissa jo product ki capabilities agents ke liye expose karta hai.
  • Human-in-the-loop: agent action se pehle insan ki approval ka intezar karta hai.
  • Human-on-the-loop: agent khud action leta hai jabke insan dekh kar zaroorat par beech mein aa sakta hai.
  • Idempotent: dohrana safe; do baar karne ka asar aik baar jaisa, taake retry hua refund do baar na jaye.
  • Provenance: maloomat kahan se ayi; agent ne asal mein kaunsi file, page ya source parha.
  • Scoped, revocable credential: aisi key jo sirf zaroori darwaze kholti aur kabhi bhi wapas li ja sakti hai.
  • Escalation: woh lamha jab agent ruk kar task insan ko deta hai kyun ke use akelay faisla nahin karna chahiye.

📚 Teaching Aid

Poori Slideshow Kholein

Poori Presentation Dekhein — Agent Experiences ki Design


Part 1 · Tabdeeli

Saada alfaaz mein: software ke khud action lene par kya badalta hai aur design ko kyun badalna parta hai.

Concept 1 · Teesra paradigm: ab aap "kaise" ki design nahin karte

Computing mein machine se baat karne ke teen tareeqe rahe hain. Batch computing mein aap poora workflow pehle batate aur intezar karte the. Command computing, yani desktop, web aur app mein, aap machine ko step by step chalate the aur goal tak kaise pohanchna hai jaanne ka bojh aap par hota tha. Har click, menu aur form aik step tha jo aapko maloom hona chahiye tha.

Teesra paradigm miqdaar mein nahin, qisam mein mukhtalif hai. Aap outcome batate hain aur agent steps chunta hai. "Kaise" ka bojh insan se machine ko chala jata hai. Isi liye agentic software obstacle course chalne ke bajaye goal tak teleport hone jaisa lagta hai.

Teen stacked cards: Batch (1945), Command (1984), aur Intent (2023 onward). Neeche track "burden of how" ko HUMAN se MACHINE ki taraf le jata hai. Batch: poora workflow pehle. Command: step by step drive aur kaise jaan'na. Intent: outcome batayein aur agent kaise sambhale. Aakhri line: pehle design tasks ko widgets mein badalti thi; ab intent, trust aur recovery ko shape karti hai.

Figure 1: Control ka markaz ulat jata hai. Machine ke "kaise" sambhalte hi designer ka kaam upstream, intent, trust aur recovery ki taraf jata hai.

Course ka har agla concept isi reversal se nikalta hai. Jab steps design nahin karte to teen nai cheezein design karte hain: insan intent kaise bayan karta hai, apne na chune steps par bharosa kaise karta hai, aur machine ghalat chune to recover kaise karta hai. Teen alfaaz yaad rakhein: intent, trust, recovery. Yeh poore course ka chota naqsha hai.

Concept 2 · Do audiences, aik system

Yeh woh baat hai jo aksar teams miss karti hain, aur ise jaldi samajhna zaroori hai kyun ke Part 3 isi par khara hai.

Aap ka agentic product aik waqt mein do qisam ke users istemaal karte hain. Insan ko samajhna aur bharosa karna hai. Doosre agents service call, data read aur insan ki taraf se action lete hain. Woh bhi users hain, bas pixels ki jagah structure parhte hain. Aage ke refund Worker ko sochiye: insan aik saada line dekhta hai, $38 refunded, undo?, jabke payment provider ka agent typed idempotent tool call dekhta hai. Idempotent yani retry safe: do calls se do refunds nahin. Aik action, do surfaces, dono durust.

Beech mein dark box "Aap ka agentic system (aik Digital FTE)". Baen human user aur daen agent user, dono two-way arrows se jure. Human ko trust surface: clarity, visible reasoning, control, consent aur undo. Agent ko machine surface: structured data, machine-readable tools, predictable contracts aur scoped provable auth. Caption: aap dono ke liye aik saath design kar rahe hain; aik ko chhorein to doosra tootega.

Figure 2: Do audiences, aik system. Baen surface trust ke liye, daen parsing ke liye. Acha product dono ko jaan boojh kar design karta hai.

Industry aik hi confusing acronym AX ke do ma'ni leti hai. Ise aik baar wazeh kar lein:

TermKis ne pesh kiyaMa'niIs course ka naam
Agentic ExperienceJohn Maedaagent ko delegate karne ka human experiencehuman surface (Part 2)
Agent ExperienceMatt Biilmann (Netlify)product ke user ke taur par agent ka experiencemachine surface (Part 3)

Dono haqeeqi aur design work hain. Zyada courses pehla sikhate hain; hum dono, kyun ke Digital FTE humans aur aas paas ke agents dono use karte hain. Unparseable machine surface par khoobsurat human surface rakhne se doosre agent ke use par system fail hoga.

Concept 3 · Interface gayab nahin hota, jagah badalta hai

Aap serious logon se suneinge ke agents interface design ka khatma hain. Users screens chhor deinge; agents browse, click aur decide kareinge; crafted UI bekaar hoga. Da'wa sanjeeda hai kyun ke kehne walon ne field banane mein hissa liya.

Lekin stakes barhne par dekhein: map app route suggest kare to bhi aap nazar dalte hain. Stakes jitne oonche, insan utna zyada verify karta hai. Har autonomous agent khamoshi se teen naye interfaces banata hai:

  • configuration surface: badalti preferences agent ko kaise sikhayein?
  • monitoring surface: zehni bojh ke baghair kaam kaise dekhein?
  • intervention surface: ghalat hone par beech mein aa kar kaise theek karein?

Imandaar position darmiyan hai. Interface gayab nahin hota; us ka center of gravity badalta hai: "task ko widget banayein" se "intent ka system shape karein". Screens ke bajaye delegated relationship ki plumbing design hoti hai: agent kya kar sakta, kaam kaise dikhata aur insan control kaise wapas leta hai.


Part 2 · Human Surface: Trust ke Liye Design

Saada alfaaz mein: product ka human-facing hissa is tarah design karna ke insan agent par bharosa, use steer aur us ki ghalti theek kar sake.

Concept 4 · Trust kamaya jata hai, maan nahin liya jata

Delegation trust par chalta hai aur agentic product mein trust sab se kam cheez hai. Insan faisla aapko de kar peechhe hota hai. Yehi lean-back poori value aur poora risk hai. Machine aik zaroori kaam mein trust torey to insan faisla wapas le kar phir nahin deta.

Trust seedha design nahin hota. Aap us ke inputs design karte hain:

Trust = waqt ke saath nazar ati reliability × check hone wali transparency × mehsoos hone wala control × undo hone wali ghaltiyan.

Yeh sum nahin, product hai: aik term ka zero sab zero. Reliable black box trust nahin kamata. Transparent magar unsteerable agent trust nahin kamata. Steer ho lekin undo na ho, tab bhi nahin. Part 2 in chaar terms ko barhata hai. Pehla qadam sab se humble: agent ko visibly uncertain rehne dein. Doubt chhupa kar confidently ghalat agent us se zyada trust torhta hai jo kahe "is hisse ka yaqeen nahin, dekh lein."

Ulta failure over-trust hai. Sau baar durust agent ke baad insan check aur drift notice karna chhor deta hai. Yeh automation complacency hai. Uncertainty visible rakhein, high-stakes work ko full autonomy na dein aur fleet-level drift dikhayein. Maqsad calibrated trust hai, maximum nahin.

Concept 5 · First contact: onboarding, cold-start trust aur har shakhs ke liye access

Concept 4 ne trust waqt ke saath dikhane ko kaha. Day one par history nahin. Har capability jump, jaise answer chatbot se acting agent, expectations reset mangti hai warna nai cheez par purani jitna trust ghalat hoga.

Teen moves first contact ko honest banate hain:

  • Har jump par expectations reset karein. Surface ko act ki power mile to wazeh kahein: "ab mein sirf bata nahin, aap ke liye kar sakta hun. Is ka matlab yeh hai." Capability badle aur mental model na badle to khatra hai.
  • Stakes kam karke trust udhaar lein. Unshown reliability ka claim na karein; lowest autonomy, action se pehle plan aur chota reversible pehla task. Day-one trust ghalti sasti karke milta hai, zero ghalti ke waade se nahin.
  • Naya hone par honest rahein. "Mein ne aap ke saath yeh pehle nahin kiya, is liye shuru mein zyada check karunga." Calibrated humility baad mein autonomy barhati hai.

Neeche aik requirement: surface har shakhs ke liye kaam kare. Agents accessibility ka khatma nahin; plain-language goal dense UI se aasaan ho sakta hai. Lekin plan, confidence, undo aur human path bina sight, mouse aur first language ke kaam karein. Part 3 ka structured labelled semantic machine surface wahi structure hai jo assistive technology parhti hai. Agent audience aur disabled users ke liye achi design aik hi simt jati hai.

WCAG 2.2 ko floor samajh kar agentic acceptance criteria testable likhein:

  • Status aur progress screen reader announce kare, sirf motion nahin.
  • Plan, undo aur human path keyboard se deep navigation ke baghair milein.
  • Confidence aur uncertainty sirf colour se na hon.
  • Insan long-running kaam pause, resume ya cancel kar sake.
  • Notifications ki quantity aur intensity adjust ho sake.
  • Har explanation ka plain-language version ho.

Concept 7 layers se jorein: Layer 1 outcome screen reader ko announce ho; Layer 3 confidence colour ke baghair; Layer 4 evidence keyboard se. Specific layer/control ke baghair criterion requirement nahin, sirf niyyat hai.

Concept 6 · Load baantein aur dikhayein kaun kya utha raha hai

Agent ka wada insan ka bojh kam karna hai: analysis aur decision ka cognitive, drafting ka creative, aur steps/coordination ka logistical load. Digital FTE banate waqt aap machine ka hissa tay karte hain.

Ghalti load ko khamoshi se move karna hai. Invisible division se agentic sludge banta hai: insan agent aur apna kaam alag nahin kar pata, safety ke liye dobara karta aur time saving khatam.

Rule: division of labor visible aur adjustable ho. Insan dekhe "yeh mein ne, yeh aap ne, yeh aap ka intezar" aur line move kare. Yeh Human-Agent Teams ke roster/role cards ka surface hai; operating model batata aur surface dikhata hai.

Udhar li vocabulary, aik line mein

Human-Agent Teams mein Worker ka role card job ki one-page spec hai: kaam, inputs, limits aur output check; roster team Workers aur un ke maqsad ki list. Yeh course un documents ka surface design karta hai.

Concept 7 · Progressive transparency: reasoning nazar aye, bojh na bane

Transparency ka trap hai. Kuch na dikhayein to black box; sab kuch dikhayein to unusable noise. Dono trust torhte hain. Progressive transparency default outcome aur zaroorat ke mutabiq reasoning depth deti hai.

Chaar stacked cards barhti depth mein. Default: outcome, agent ne kya kiya aik saada line. Phir plan, ordered steps. Phir why, confidence signal ke saath. Sab se gehra evidence, sources, tool calls aur full trace. Baen downward arrow "depth".

Figure 3: Progressive transparency. Default aik honest line; har depth aik tap neeche, reader par force nahin.

Chaar layers:

  1. Outcome, aik plain line: agent ne kya kiya.
  2. Plan: ordered steps, kaam ki shape check karne ke liye.
  3. Why: rationale aur honest confidence signal. Fake percent nahin, asli high / low / unsure. "73%" jhooti calibration, "unsure" sach.
  4. Evidence: sources, tool calls, full trace; audit aur debugging ke liye.

Chote surfaces kaam karte hain: provenance chip ("aapki 3 files par mabni"), shaky hissa uncertainty marker, aur trace link. Model mathematics nahin; human ka sawal: kya bharosa karun, aur nahin to pehle kahan dekhun?

Concept 8 · Autonomy dial: barhti hui permission

Autonomy "off" se "sab kuch" ka switch nahin. Yeh dial hai jo insan, aap nahin, hold kare. Low ship karein aur agent ke prove hone par barhne dein, jaise manager trust ke baad new hire ko azadi deta hai.

Paanch-stop dial. 1 Suggest: agent propose, aap act. 2 Confirm: har step se pehle poochhe. 3 Act in limits: budget ke andar act, bahar poochhe. 4 Act and report: act phir report. 5 Autonomous: akela run, aap audit. Stops 1–2 human in the loop; 3–5 human on the loop. Gold arrow: verified reliability se trust barhta hai.

Figure 4: Autonomy dial. Human in the loop mein agent rukta hai; human on the loop mein act karta aur aap supervise karte hain. Reliability dial barhati hai.

Do ideas safety dete hain. Human-in-the-loop action se pehle approval ka intezar; human-on-the-loop action aur human intervention. Low-stakes, reversible, proven kaam on-the-loop kamata hai; high-stakes ya irreversible kaam reliability ke bawajood in-the-loop. Doosra, per-task consent ke saath safe defaults: naya agent "suggest" se shuru, har stop insan jaan boojh kar chunta hai.

Yeh Nervous System approval gates aur Digital FTE ke authority model ka front-of-house hai. Backstage approval durable audited event; front par mehsoos hone wala dial.

Concept 9 · Intent preview aur plan-review aadat

Sab se sasti ghalti woh hai jo abhi nahin hui. Action se pehle, khaas irreversible par, plan dikhayein aur edit karne dein: "mein 1, 2, 3 karne wala hun. Kuch badlein?" Yeh intent preview baad ki explanations se zyada regret rokta hai.

Cowork ka plan review yehi pattern hai; yahan ise ship hone wale product mein banate hain:

  • Preview stakes ke saath scale kare. Aik internal draft par quiet "bhej raha hun, undo?"; 500 emails ya money par full plan aur explicit confirm.
  • Plan mid-flight editable ho, sirf approvable nahin. Accept/reject wall hai; "haan, step 2 chhor dein" partnership.

Concept 10 · Asynchrony ke liye design: kaam ki nai rhythm

Command software synchronous tha. Agentic work mein intent set, disconnect, aur progress ya done work par wapsi hoti hai. Is collaboration shape ko apni design chahiye.

Paanch-node loop: Set intent → Agent works → Nudge → Review → Refine → start. Human nodes terracotta, agent-alone slate. Center: You can disconnect. Legend human present aur agent alone.

Figure 5: Asynchronous loop. Insan intent aur review par; darmiyan agent akela, sirf real decision par interrupt.

Chaar surfaces:

  • Intent capture aap ke baghair chalne jitna saaf.
  • Glanceable progress: teen seconds ka status, logs ki deewar nahin.
  • Nudge, notify nahin. Sirf real decision par interrupt. Har step ping needy coworker hai. Interruption earn ho.
  • Wapsi ka review-and-refine surface, jahan finished work inspect, correct aur re-aim ho.

Waiting ka felt experience bhi design karein. Agent slow aur mehnga ho sakta hai. Silence broken lagti hai, busy nahin. Current step ka honest progress, upfront rough time, expensive runs ka cost/budget aur kaam roke baghair check-in dikhayein. Latency aur cost hidden backend nahin, experience hain.

Concept 11 · Repair and redress: ghalat din ke liye design

Agent ghalat action lega. Probabilistic system real work mein kabhi fail hoga. Recoverability end par bolted error state nahin, shuru se first-class surface hai.

Chaar moves betrayal ko bump banate hain:

  1. Task jitna de, undo aik click ke qareeb. Sab se strong trust-builder; log reversible agent ko autonomy dete hain.
  2. Seedhi maafi aur plain account, hedging ya user blame nahin.
  3. Corrective action aur next step: "transfer reverse, review flag."
  4. Human tak visible path, hamesha; accountability aur de-escalation.

Do metrics: escalation frequency ka practitioner starting band 5–15%; kam par guessing, zyada par timidity. Recovery success 90% se oopar chahiye. Domain ke mutabiq calibrate karein, laws nahin. Yeh ops aur UX dono hain.

Concept 12 · Kai agents ki supervision: aik se workforce

Das Workers par koi das plans, approvals ya traces nahin dekh sakta. Design monitoring se exceptions ki triage ban jati hai.

Supervisor fleet view. "Needs you now" mein $900 dispute Refunds Worker aur doubled escalation se drifting Support Worker. Neeche "Running clean — in the log" mein Invoicing, Onboarding, Research, Billing. Footer: 6 Workers, 2 need you, 4 log mein; attention budget.

Figure 6: Fleet view. Surface human attention un chand Workers par kharch karta hai jinhein insan chahiye, baqi log mein.

Teen surfaces:

  • Fleet view: aik nazar mein running, blocked, waiting; operations board, das chats nahin.
  • Attention triage: $900 dispute oopar, 200 clean refunds log. Surfaced:silent ratio attention budget.
  • Drift, sirf distress nahin: "escalation rate doubled" loud failure se pehle. Surface human ke liye over-trust pakarta hai.

Drift response trigger kare. Fleet-level circuit breaker Worker ke apne baseline multiple par autonomy kam ya pause karke review uthaye. Threshold magic number nahin, calibrate hota hai. Default slipping Worker ko kam autonomy.

Yeh Human-Agent Teams aur Paperclip roster/control plane ka front-of-house hai: woh team define karta, yeh room design karta jahan insan team dekhta aur tay karta hai kya nahin dekhna.


Part 3 · Machine Surface: Agents ko Users Samajh Kar Design

Saada alfaaz mein: product ka woh hissa design karein jo doosre agents use karte hain, taake pehli koshish mein durust use ho.

Concept 13 · Agent Experience (AX): product ke robot users

Ab woh half jo teams design nahin kartin. Product ko agents use karte hain: users ki taraf se acting agents aur apni workforce Workers. Woh layout nahin, structure parhte hain. Hostile structure mein human ka agent silently fail aur human aapko blame karta hai.

"Machine surface" ke neeche chaar cards. Access: agent kis ki authority? Scoped, revocable. Context: model meaning samjhe? Tools: machine-readable typed discoverable? Orchestration: safe chaining, contracts, idempotency, limits? Caption: well-designed connector/MCP server acha AX aur agents ka interface.

Figure 7: Machine surface. Chaar sawal decide karte hain agent product use kar sakta hai; har aik design decision.

  • Access: scoped revocable credential se agent kis ki authority prove kare? AI Identity ka masla.
  • Context: model product ka ma'ni samjhe? Clear names, honest descriptions, readable semantics.
  • Tools: capabilities machine-readable, typed, discoverable hain ya scrape wali UI mein?
  • Orchestration: predictable contracts, idempotent actions aur sane limits se safe chaining?

Well-designed connector ya MCP server acha AX hai. Skills & Connectors aur Connector-Native Apps machine-reader interface design hain. SKILL.md aur typed MCP tool robot-user UX hain.

Trustworthy machine surface ki habits:

  • Tool action par plainly name: refund_order, process nahin.
  • Schemas narrow, typed, validated.
  • Side effects/danger declare, dangerous tools confirm ya policy check.
  • Structured actionable errors: code, next step, retryable, retry_after, fallback.
  • Jahan mumkin actions idempotent, retry double-charge/send na kare.
  • Provenance aur permission: data source, authority, revocation.
  • Docs agent ke liye: examples, limits, failure modes, retry.
  • Contract test, sirf screen nahin.

2026 tak MCP tool discovery/call, resource read aur access authentication standardize karta hai. MCP Apps UI metadata deta hai: tool ui:// interface ko _meta.ui.resourceUri se link karke widget la sakta hai. Magar orchestration, governance aur long-running state aap ke hain. Protocol agent ko darwaze se andar lata hai; andar authority aur supervision design hai.

Good AX ke do halves: protocol conform, phir omitted policy, server allowlists, consent gates, spend limits, audited logs. MCP Apps mein sandboxing/controls host ka kaam hain. Wire format commodity, trust policy aapki.

Concept 14 · Generative UI: interfaces lautane wale agents

Do audiences aik surface mein milte hain. Tool sab hosts ko normal text aur Apps hosts ko interactive ui:// resource de sakta hai, jo _meta.ui.resourceUri mein named hai. Host conversation ke sandbox iframe mein render karta hai.

"Generative UI with MCP Apps" flow. Tool ka _meta.ui.resourceUri ui:// interface, text fallback. Host Claude, VS Code, Goose HTML fetch. Widget sandboxed: no cookies, host page, escape. Feedback JSON-RPC over postMessage. Callout: safe like data, expressive like code; one widget many hosts.

Figure 8: Tool interface declare, host conversation sandbox mein render, aik audited channel wapas; text har jagah fallback.

Generative UI random page ya arbitrary injected code nahin, tool call se attached task interface hai. Tool fallback text aur interface declare; capable host widget render; warna text decision. Principle "safe like data, expressive like code": sandbox content, loose client code nahin.

Aik capability ke do surfaces: agents ka MCP contract aur humans ka widget. Aik call approval card, dashboard, map ya chart aur parseable tool dono de sakta hai.

Teen lines mein positioning

MCP machine surface hai. MCP Apps interactive human surface. Dono aik tool ko humans aur agents dono ke liye banate hain.

MCP Apps November 2025 proposed aur July 2026 finalized first official UI extension hai, MCP-UI, OpenAI aur Anthropic ka joint kaam. Claude, Desktop, VS Code, Goose, Postman render karte; OpenAI Apps SDK isi base par ChatGPT apps. January 2026 Claude launch mein Asana, Slack, Figma, Canva, Box, Hex interactive connectors production mein the. One widget, many hosts, features vary.

Chaar experience properties:

  • Context rehta hai: app conversation mein, no tab switch.
  • Dono taraf baat: widget server tools call, host results push; API/login/state protocol se.
  • Consent se host powers: outcome host ke connected services se route hota hai.
  • Construction se safe: sandbox host page/cookies/storage rokta, messages audited.

Widget complex data, many-option configuration, rich media, real-time monitoring aur multi-step workflow ke liye. Plain text kaafi ho to text. Widget bhi nudge ki tarah jagah earn kare.

Portability: progressive enhancement. Open standard first; payment/store extras feature-detect aur baqi hosts par gracefully degrade. Vendor-only surface phans jata hai.

Chat/app wall toot ti hai. Agent task ka exact form/chart/map compose, styling/security/components aap ke control. Ab fixed screens nahin, agent ke bolne layak component vocabulary design hoti hai.

Ubhar raha aur badal raha hai

MCP Apps July 2026 spec mein finalized magar active development mein hai. Pattern ke shape ke liye design karein: data ki surat interface, inescapable sandbox. Build se pehle modelcontextprotocol.io/extensions/apps confirm karein. Appendix A anatomy aur Lab build deta hai.


Part 4 · Nai Craft

Saada alfaaz mein: job ki nai skills, safety, measurement aur har Worker ka one-page document.

Concept 15 · Naye design objects aur choreographer ka kaam

Screens nahin to kya draw? Nai primary objects:

  • Policy surfaces: permissions, spend ceilings aur ethical boundaries. Rules aur unhein set karne ke controls.
  • Confidence conveyors: Concepts 7/11 ke provenance chips, uncertainty markers aur clean rollbacks; system sure hone ka sach kaise bataye.
  • System temperament: agent kitna patient/proactive, kitni baar aur kaise bole; brainstorm mein eager, money par cautious.

Role screen-crafter se choreographer hai: humans aur agents ka saath movement, information architecture, conversation, operations aur control kab rakhna/hatana. Choosing Agentic Architectures ka backstage pattern, single agent/planner/multi-agent, front par human supervision tay karta hai. Architecture aur experience aik decision ke do views.

Concept 16 · Surface safety control hai

Surface trust banata aur bachata hai. World-acting agent ka attack surface hai aur defense ka bara hissa design hai. Shared list OWASP Top 10 for LLM Applications hai.

Do-column "surface as safety control" map. Agent risks: prompt injection LLM01, excessive agency LLM06, misinformation/over-trust LLM09, unbounded consumption LLM10, sensitive-data disclosure LLM02. Responses: provenance+intent preview; autonomy dial+policy surfaces; visible uncertainty; cost meter+spend limits; scoped revocable access. Footer: safety control, sirf display nahin.

Figure 9: Har OWASP risk ko course ka design response; safety experiential feature hai, backend chore nahin.

  • Prompt injection LLM01: hidden web/document instruction. Provenance read source dikhata aur intent preview consequential action check karta hai.
  • Excessive agency LLM06: least authority default, high stakes in-loop, capabilities intentional.
  • Misinformation/over-trust LLM09: visible uncertainty aur provenance shaky claim ko sure jaisa nahin dikhate.
  • Unbounded consumption LLM10: runaway loop/denial-of-wallet par visible cost aur spend limits.
  • Sensitive-data disclosure LLM02: scoped revocable access aur visible withdrawable consent.

Do rules. Surface safety control hai, sirf display nahin; clutter ke naam par provenance, uncertainty, preview ya cost hatana protection hatana hai. Har system ko governance surface chahiye: capability, permissions, incident, audit aur sab se zaroori, aik move mein agent pause/kill ka owner.

Governance surfaceDesign sawal
Capability approvalWorker capability kaun barha sakta hai?
Permission reviewTool scopes/data access kaun approve?
Incident reviewFailure/postmortem ka owner?
Audit logActions, plans, calls, approvals kaun dekhe?
Kill switchAik move mein pause/disable/rollback kaun?
Drift reviewRising escalation ya slipping recovery kaun investigate?

NIST AI Risk Management Framework in organizational controls ko Govern, Map, Measure, Manage kehta hai; Human-Agent Teams aur Workforce with Paperclip operationalize karte hain. Surface par control reachable banayein, policy invent nahin.

Concept 17 · Experience measure karna

Elegant surface fail ho sakta hai. Eval-Driven Development Worker output correct hai ya nahin, unit/tool/trace/safety/regression evals se measure karta hai. Experience metrics relationship: insan delegate, trust, steer aur recover kar sake. Correct Worker bhi unsupervisable surface rakh sakta hai.

MetricKya batata hai
Plan-acceptance rateSamajh kar approve ya blind rubber-stamp/reject?
Intervention rateHuman kitni baar aya? Falling trend trust.
Recovery successFailure/escalation ke baad acha end?
Over- vs under-trustBad accept ya good reject?
Notification precisionKitne interruptions worth the?
Time saved vs attention spentTotal burden kam ya sirf move?

Trends, snapshots nahin; last month ke baghair 12% be-ma'ni. Microsoft HAX Playbook ki tarah launch se pehle failures rehearse aur recovery design. Sab se aham time saved versus attention spent; negative ho to surface fail.

Launch se pehle tests:

  1. Plan-review test: action se pehle read, understand, correct.
  2. Over-trust test: subtle wrong output par uncertainty blind approval roke.
  3. Recovery test: reversible ghalti aur clean undo ka waqt.
  4. Interruption test: useful nudges measure.
  5. Accessibility pass: sirf keyboard/screen reader se plan, progress, undo, escalation.
  6. Machine-surface contract test: wrong type reject, danger gated, error structured, retry idempotent.

Pass correctness proof nahin, woh Eval-Driven Development hai; magar ghalti par experience qaim rehne ka proof hai.

Concept 18 · Anti-patterns aur aakhri design brief

Anti-patternKaisaToota concept
Black boxreasoning baghair action7 · progressive transparency
Agentic sludgesab re-check, no time saving6 · visible division
Over-eager agentday one high autonomy8 · autonomy dial
Over-trusted agentunchecked work par high autonomy4 · calibrated trust
Notification spamhar step ping10 · nudge, don't notify
Alarm-fatigued consolehar Worker ping, human ignore12 · attention triage
False confidenceshaky guess sure fact4 · visible uncertainty
Trap doorundo nahin11 · repair/redress
Confused deputyuser aur injection alag nahin16 · safety surface
Unmeasured surfaceelegant, help unmeasured17 · measurement
Mystery-meat APIagent parse na kare13 · Agent Experience

Deliverable: pehle Digital FTE ka Agent Experience Brief, gyarah decisions:

  1. Do audiences: human aur agent users.
  2. First contact: onboarding, capability reset, sight/mouse baghair.
  3. Trust surface: default aur neeche teen layers.
  4. Load map: human/Worker division aur visible line.
  5. Autonomy ladder: paanch stops, forever in-loop actions.
  6. Async plan: intent, progress, wait, nudges.
  7. Recovery plan: undo, escalation, do health metrics.
  8. At scale: fleet, attention, drift.
  9. Machine surface: connector/MCP tools, AX pillars.
  10. Safety surface: agent threats aur defenses.
  11. Scorecard: experience metrics aur correctness boundary.

Yeh artifact Worker experience ke liye Human-Agent Teams operating docs jaisa hai. Spec-Driven Development mein experience layer ki spec: non-deterministic Worker ka deterministic reviewable surface. Appendix C blank version.


Worked Example: Support Worker ke Do Surfaces

Digital FTE course ka customer-support FTE, dono surfaces.

Human surface. Console mein har ticket aik line: "Refunded order #4021, $38, confidence: high." Layer 1. Tap par plan: order, policy, refund, email. Phir why/policy clause. $38 limit mein, stop 3 act within limits. $900 chargeback limit se bahar, stop 2 in-loop. Har refund 24-hour undo. Lead on the loop, glance karta hai. Das Workers mein fleet $900 dispute uthata, clean refunds log mein.

Machine surface. Company orchestrator isi Worker ko call karta aur payment provider ke liye yeh agent hai. Typed hard-limit MCP tool (tools), scoped revocable merchant credential (access), valid refund ki clear description (context), idempotent refund (orchestration). Human surface par hidden magar system in ke baghair collapse.

Aik Worker. Do audiences. Do jaan boojh kar designed surfaces.

Stakes barhein. Refund ki jagah irreversible vendor payments. Undo safety net nahin, design recovery se prevention: mandatory detailed intent preview. Autonomy "act within limits" se oopar nahin; low limit, large payment forever in-loop. Confidence bar barhta. Provenance aur second human approver lazmi. Governance: kill switch, full audit, threshold payments ka named owner. Same patterns sakht, kyun ke ghalti mehngi. Recovery/prevention dial payroll, clinical triage, grading aur regulated/irreversible domains ke liye hai.


Hands-On Lab: Apna Pehla MCP App Banayein

Part 3 ke machine-surface design ko ab real MCP App mein banayein: text wall ke bajaye working interactive widget jo Claude ya supporting host ke andar render ho.

Kitab ki tarah coding agent ko direct karke build karein. Official guide ke mutabiq AI coding agent aur MCP Apps skill fastest tareeqa hain. Skill architecture/best practices, agent typing aur aap spec/judgment dete hain.

Zaroorat. Node.js 18+, terminal, Skills-supporting agent: Claude Code, OpenCode, Codex, Cursor, Gemini CLI, Goose. Claude custom-connector test paid plan; Step 3 local host free.

Step 1 · create-mcp-app skill install karein

Skills & Connectors ke mutabiq skill instructions/examples folder hai. Official create-mcp-app architecture aur pitfalls sikhata hai.

Claude Code plugin:

/plugin marketplace add modelcontextprotocol/ext-apps
/plugin install mcp-apps@modelcontextprotocol-ext-apps

Doosre agents:

npx skills add modelcontextprotocol/ext-apps

Manual: github.com/modelcontextprotocol/ext-apps clone aur plugins/mcp-apps/skills/create-mcp-app ko ~/.claude/skills/, ~/.codex/skills/, ya ~/.cursor/skills/ mein copy.

Verify:

What skills do you have access to?

List mein create-mcp-app ho to agent MCP Apps banana janta hai.

Step 2 · Das-minute loop: scaffold, build, serve

Agent ko aik line:

Create an MCP App that displays a color picker

Agent skill load karke MCP server, widget UI aur build config scaffolds karega. Project folder mein:

npm install && npm run build && npm run serve

Server http://localhost:3001/mcp par run hai. Ab render dekhein.

Step 3 · Render dekhein

Option A: free local host. ext-apps minimal host:

git clone https://github.com/modelcontextprotocol/ext-apps.git
cd ext-apps/examples/basic-host && npm install
SERVERS='["http://localhost:3001/mcp"]' npm start

http://localhost:8080 kholein, tool call aur sandbox iframe widget dekhein.

Option B: Claude web/Desktop. Doosre terminal mein tunnel:

npx cloudflared tunnel --url http://localhost:3001

Generated https://….trycloudflare.com ko Settings → Connectors → Add custom connector mein add karein. Pro/Max/Team plan chahiye. Nai chat mein color picker maangein.

Step 4 · Agent ka build parhein

Do MCP primitives, aik bridge. Server par UI metadata wala tool aur UI serve karta resource:

// server.ts (the load-bearing lines)
const resourceUri = "ui://get-time/mcp-app.html"; // ui:// marks this as an App interface

registerAppTool(
server,
"get-time",
{
title: "Get Time",
description: "Returns the current server time.",
inputSchema: {},
_meta: { ui: { resourceUri } }, // the one line that turns a tool into an App
},
async () => ({
content: [{ type: "text", text: new Date().toISOString() }], // the text fallback
}),
);

registerAppResource(
server,
resourceUri,
resourceUri,
{ mimeType: RESOURCE_MIME_TYPE },
async () => ({
contents: [{ uri: resourceUri, mimeType: RESOURCE_MIME_TYPE, text: html }],
}),
);

Widget ka App class only sandbox channel:

// src/mcp-app.ts (the load-bearing lines)
const app = new App({ name: "Get Time App", version: "1.0.0" });
app.connect(); // open the postMessage channel to the host

app.ontoolresult = (result) => {
/* the host pushes the first tool result here */
};

await app.callServerTool({ name: "get-time", arguments: {} }); // the UI calls tools back

Text content non-Apps fallback. Widget host page/cookies nahin chhoota; postMessage par audited JSON-RPC. Har callServerTool round trip, is liye wait gracefully.

Step 5 · Asal build: refund approval card

Vibe nahin, spec; widget par Spec-Driven Development:

Using the create-mcp-app skill, build an MCP App called refund-approval.

Tool: review_refund(order_id: string, amount: number, confidence: "high" | "low" | "unsure").
It returns the refund details as plain text (the fallback) and renders an approval card.

The card must:
1. Show one plain line: "Refund #<order_id> · $<amount> · confidence: <word>".
Confidence is always a word, never a colour.
2. Offer two buttons, Approve and Escalate to a human. Both must be reachable
by keyboard, with labels a screen reader announces.
3. On Approve, call the server tool approve_refund(order_id), then show
"Approved · Undo available for 24h" with an Undo button that calls
undo_refund(order_id).
4. If amount > 50, disable Approve and show "Above limit: needs a human",
leaving only Escalate active.
5. Make approve_refund and undo_refund idempotent on order_id: calling either
twice must be safe.

Steps 2/3 jaisa build, serve, test, phir design pass:

CheckConcept
Non-Apps host ya text content: fallback decision rakhta hai?14 · fallback
Confidence word, colour nahin; plain line?7 · transparency, 5 · accessibility
Approve phir Undo, one click aur twice safe?11 · repair, 13 · idempotency
$900 refund refuse aur human route?8 · autonomy dial
Keyboard se har control?5 · access
Action names, review_refund, process nahin?13 · machine surface

Sab pass par aik call ne human plain line aur agent typed idempotent contract dono diye: Concept 2 ke dono surfaces.

Step 6 · Mazeed gehrai

Source of truth: modelcontextprotocol.io/extensions/apps/overview, modelcontextprotocol.io/extensions/apps/build, apps.extensions.modelcontextprotocol.io, aur ext-apps GitHub examples: maps, 3D, PDFs, dashboards, React/Vue/Svelte/vanilla starters. Command fail ho to agent guide fetch karke reconcile kare.

Verified, magar details ko halka pakrein

Mid-2026 guide se commands/patterns verified. July 2026 finalized extension active development mein; package names/helpers/hosts badleinge. Durable pattern tool + ui:// resource + sandbox render + postMessage hai.


Projects

  1. Used agent audit. Concept 18 anti-pattern score, weakest trust input aur best change.
  2. Dial draw. Worker ke five stops, forever in-loop actions aur wajah.
  3. Nudge budget. Interrupt events se sirf truly human-needed list.
  4. Machine surface likhein. Capability connector/MCP tool unfamiliar agent ke first-try use ke liye, chaar AX sawal.
  5. Fleet view. Paanch Workers, glance, interruption/log aur drift signal.
  6. Full brief capstone. Digital FTE ka eleven-section Agent Experience Brief.
  7. Widget ship. Lab Step 5 tak real host, design pass aur plain form se mukhtalif teen decisions.

Reader track

Build ke bajaye direct karna ho to Concepts 1–4, 8, 12, 13, 16, 17; phir autonomy ladder, fleet view aur machine-surface judgment. Leader ke trustworthy review ke liye kaafi.

Kitab Mein Is ki Jagah

Yeh Spec-Driven Development jaisi design discipline hai, install tool nahin. Human-Agent Teams operating model likhta; yeh insan ka team-work surface. Dono mil kar manual aur control room.

Appendix A: MCP Apps ka Naqsha (2026)

Concepts 13/14 aur Lab ke parts, mid-2026 status. July 2026 finalized magar active; modelcontextprotocol.io/extensions/apps confirm karein. MCP Apps tool layer par hai jahan discovery, calls, results aur task UI hoti hai.

PieceKya hai
Toolnormal MCP tool, _meta.ui.resourceUri interface; text non-Apps fallback
ui:// resourceHTML interface, aam tor par CSS/JS bundled, server resource
Sandboxed iframeisolated render, host page/cookies/storage se band
postMessage channelJSON-RPC, ui/ methods aur tools/call, host-auditable
csp aur permissionsexternal origins aur camera/mic capabilities
App class@modelcontextprotocol/ext-apps: connect(), ontoolresult, callServerTool(); optional web-API wrapper
Host supportmid-2026 Claude, Desktop, VS Code, Goose, Postman, MCPJam; OpenAI Apps SDK same base

Tool/resource machine, widget human, sandbox/channel/permissions safety surface. Aik pattern, teen surfaces.

Appendix B: MCP Apps vs OpenAI Apps SDK, Designer Note

Do common paths MCP Apps aur OpenAI Apps SDK rivals nahin. OpenAI Apps SDK MCP par built hai: ChatGPT App MCP server with extras; same iframe, JSON-RPC aur UI declaration.

Experience extras:

  • Discovery surface. ChatGPT store aur Claude directory (claude.ai/directory); find aur pre-use trust ecosystem cold-start.
  • In-chat payment. 2026 beta/selected markets; review-confirm-pay compressed, is liye clear commitment aur reversible recovery.
  • Distribution. User-base reach, single-host lock-in tradeoff.

Rule: open MCP Apps base first, vendor store/checkout extras feature-detect aur baqi jagah gracefully degrade. Vendor-only phansata hai.

Build detail: Payment-Enabled Agents, Connector-Native Apps, Plugins for AI Agents. Vendor details current docs se confirm.

Appendix C: Agent Experience Brief (Fillable Template)

Concept 18 ka blank deliverable. Worker name aur har field bharein. Mushkil field pending design decision hai. Specific questions ki wajah se yeh generation prompt ya Agent Factory spec bhi hai. Aik-do pages tak rakhein.


Agent Experience Brief: Worker name: ____________________________ · Owner: __________________ · Date: __________

1 · Do audiences · Concept 2 Human aur agent users kaun? ____________________________________________________________________

2 · First contact · Concept 5 Onboarding, action-power expectation reset aur sight/mouse baghair WCAG 2.2 use? ____________________________________________________________________

3 · Trust surface · Concept 7 Default Layer 1 aur plan, why + confidence, evidence? ____________________________________________________________________

4 · Load map · Concept 6 Human/Worker kaam aur visible movable line? ____________________________________________________________________

5 · Autonomy ladder · Concept 8 Five stops aur forever human-in-loop actions? ____________________________________________________________________

6 · Async plan · Concept 10 Intent, progress, wait latency/cost aur nudges? ____________________________________________________________________

7 · Recovery plan · Concept 11 Undo, escalation aur do health metrics? ____________________________________________________________________

8 · At scale · Concept 12 Fleet view, attention aur drift? ____________________________________________________________________

9 · Machine surface · Concept 13 Connector/MCP tools, AX access/context/tools/orchestration? ____________________________________________________________________

10 · Safety surface · Concept 16 Injection, excessive agency, runaway cost, disclosure defenses? ____________________________________________________________________

11 · Scorecard · Concept 17 Experience metrics aur correctness handoff? ____________________________________________________________________


Appendix D: Bhara Hua Brief (Worked Example)

Customer-support Refund Worker template; har jawab concrete decision, sawal ki repetition nahin.

Agent Experience Brief. Worker name: Refund Worker · Owner: Support Lead · Date: 2026-07-01

1 · Do audiences. Human support lead; agents ticket orchestrator aur payment API.

2 · First contact. "Suggest" ship; act promotion banner: "ab $50 tak refund khud issue, sirf recommend nahin." Controls keyboard/screen reader; confidence word, colour nahin.

3 · Trust surface. Layer 1 line; Layer 2 order→policy→refund→email plan; Layer 3 why/policy/high-low-unsure; Layer 4 trace/raw API.

4 · Load map. Worker logistics; lead above-limit/low-confidence judgment; "waiting on you" lane.

5 · Autonomy ladder. 1 Suggest → 2 Confirm → 3 Act within limits ≤ $50 → 4 Act/report → 5 Autonomous. Stop 3; >$50, chargeback, fraud flag forever in-loop.

6 · Async plan. Aik ticket/goal, named progress; no cost meter; nudges above limit, low-confidence match, provider error.

7 · Recovery plan. 24-hour one-click reversal; lead then on-call finance; escalation 5–15%, recovery >90%.

8 · At scale. Ten Workers; $900 aur doubled escalation surfaced; reversal >2x own baseline par stop 2.

9 · Machine surface. Scoped revocable refund credential; valid-use context; typed refund_order(order_id, amount, reason) $50 cap; order_id idempotent.

10 · Safety surface. Ticket data, plan provenance; $50 cap/in-loop pins; bounded spend; refund-only credential.

11 · Scorecard. Plan acceptance, intervention trend, recovery, time-vs-attention. Decision correctness Eval-Driven Development ka kaam.


Sources aur Mazeed Parhai

  • John Maeda, Simplicity and Agentic Experience aur Design in Tech Report 2026: human surface.
  • Matt Biilmann, agents-as-users ke Access, Context, Tools, Orchestration.
  • Microsoft Design, Space/Time/Core, nudge aur calibrated uncertainty.
  • Adrian Levy, collaboration, load, transparency, async, dual audiences.
  • Smashing Magazine, intent preview, autonomy, intervention, repair benchmarks.
  • Jakob Nielsen, third UI, No More UI, accessibility provocation aur rebuttals.
  • MCP Apps SEP-1865 aur Model Context Protocol: tool ui:// resource ko _meta.ui.resourceUri se link; sandbox iframe aur JSON-RPC. Overview modelcontextprotocol.io/extensions/apps/overview, build modelcontextprotocol.io/extensions/apps/build, SEP modelcontextprotocol.io/seps/1865-mcp-apps-interactive-user-interfaces-for-mcp, API apps.extensions.modelcontextprotocol.io, ext-apps repository, create-mcp-app skill aur Claude launch claude.com/blog/interactive-tools-in-claude. Host sandbox/allowlist/consent/audit deployer ka kaam.
  • OWASP Top 10 for LLM Applications 2025, Concept 16 threats.
  • NIST AI RMF 1.0, Govern/Map/Measure/Manage.
  • Microsoft HAX Toolkit/Playbook, pre-launch failure rehearsal.
  • W3C WCAG 2.2, accessibility floor.

Flashcards Study Aid


Apni Samajh Test Karein

Checking access...