What an AWS SA does all day, how to run a working backwards session, the talk you will give in the loop, and what to say when a customer pushes back.
About 50 minutes to read. The quick version takes 3. Chapter 3 (your talk) is the one to rehearse out loud.
The quick version
An SA is the customer's trusted technical advisor. You design with them, prototype to prove a point, and pull in the right AWS help. You do not write their production system.
Bill's line is the brief: "show our customers how to build the next product that will delight their customers." That is Working Backwards plus a fast path to a working prototype.
Your talk: Ask the Declaration. A live, public, cited GenAI app with a real eval and an abstain threshold. Then show the same design on AWS for Beacon. Backup: PROBE.
In every customer conversation: ask before you answer, name the trade-off, offer one concrete next step.
Objections are almost always about trust, cost, or skills. Answer with a mechanism (eval set, human review, cost per task, EBA), never with a slogan.
Saying no is part of the job. Say it with a reason and a route to yes.
What an AWS SA does
Start here, because half the loop checks whether you understand the job. Many strong engineers fail it by sounding like a consultant who will build the system, or a salesperson who will close the deal. The SA is neither.
The trusted technical advisor
The problem
Beacon's CEO wants GenAI features in two quarters. Beacon has 250 engineers, none of whom have shipped a GenAI feature to 911 centers. If they guess, they burn a quarter on a demo that legal kills, or they pick an architecture that GovCloud cannot host.
Picture it
A climbing guide on a hard mountain. The guide does not carry you up. You climb. The guide has been on this face many times, knows which route has loose rock, can call in a rescue team or a gear supplier, and tells you straight when the weather says turn back. You trust the guide because they would rather lose the summit than lose you.
In plain words
You help the customer's engineers and leaders make good technical decisions on AWS, faster than they could alone. You whiteboard designs, review architectures, build small prototypes to prove a point, teach, and connect them to the right people inside AWS. You bring their needs back to AWS service teams. The customer owns and builds the product.
The real term
AWS describes SAs as wearing three hats: technical advisor (Well-Architected reviews, whiteboarding), customer advocate (representing the customer inside AWS, including filing product feature requests), and educator (Immersion Days, workshops). The SA works inside an account team with an account manager. The goal is a technical win: the customer chooses AWS for a workload because the design works and they trust it.
At Beacon
You meet the CTO and VP Engineering, run a working backwards session on Dispatch Assist, whiteboard a Bedrock design that works in GovCloud, bring in a PACE prototyping team or run an EBA, and file a feature request when a model they need is missing in GovCloud. Beacon's engineers ship it.
Go deeper (for follow-up questions)
SAs are free to the customer. That is why the role runs on trust: nobody pays you, so the customer listens only if your advice is right and in their interest.
The line an interviewer probes: "What is the difference between helping and doing their job?" Your answer: I build to prove a decision (a prototype, a reference design, a spike), and I hand it over with the reasoning. I do not become the owner of their production code, because then they cannot run it at 2 AM when a 911 center calls.
Who does what around the customer
The interviewers will ask "who else would you bring in?" This table is the answer.
Role
What they do
At Beacon, you would call them when
Solutions Architect (you)
Technical advisor across the whole account. Breadth, design, trust, prototypes to prove a point.
Always. You are the technical lead for the relationship.
Account manager
Owns the commercial relationship: pricing, agreements, credits, executive meetings, the business plan.
The CEO asks about cost commitments, credits, or an executive sponsor.
Specialist SA
Deep expert in one domain (GenAI, security, databases, analytics). Works across many accounts.
You need depth on Bedrock Agents, OpenSearch tuning, or CJIS controls beyond your own.
AWS Professional Services (ProServe)
Paid consulting delivery. They build or migrate with the customer's team under a contract.
Beacon wants build capacity, not advice, and will pay for it.
AWS Partner (consulting)
Third-party firms that deliver projects, often funded in part by AWS programs.
Same as ProServe, often cheaper or more specialized in public sector.
PACE (prototyping team)
AWS builders who make working GenAI and agentic prototypes with a customer in weeks.
A strategic feature needs a real prototype fast and qualifies for their time.
Generative AI Innovation Center
AWS scientists and strategists who take high-value GenAI ideas from use case to production.
The problem needs science depth (custom evals, fine-tuning, agent design) and has a strong business case.
Service teams
The engineers and product managers who build Bedrock, OpenSearch, and so on.
A missing feature blocks Beacon. You file and champion a product feature request.
The SA sits between the customer and the rest of AWS, and routes each need to the right team.
What "Product Acceleration" means for an ISV
The problem
An ISV (independent software vendor) like Beacon is not buying AWS to run its own back office. AWS is inside the product it sells. When Beacon ships a feature slowly, 300 cities and 150 districts wait, and a competitor may get there first.
Picture it
A restaurant supplier who also helps the chef design the new menu. The supplier sells more ingredients only if the new dishes are good and customers come back. So the supplier's chef-in-residence spends time in the kitchen, tasting and suggesting, not only taking orders.
In plain words
You help the ISV get new product features from idea to customers faster. That means helping them pick the right feature (working backwards), prove it works (prototype, EBA), build it on a sound design (architecture reviews), and get it to market (Marketplace, co-sell).
The real term
AWS's public sector blog describes a Product Acceleration Team that "helps customers iterate and bring new solutions to market using Amazon's Working Backwards process." The GovTech ISV SA posting names the work: discovery, GenAI on Bedrock with agentic patterns and RAG, moving prototypes to production under public sector rules (CJIS, StateRAMP, FedRAMP, HIPAA), modernization, and AI-assisted development tools such as Kiro and Amazon Q Developer.
At Beacon
A quarterly rhythm: one working backwards session per new AI feature, one architecture review before each build, one EBA or PACE engagement for the riskiest feature, and a Marketplace listing so districts and counties can buy through existing AWS contracts.
How SAs are measured
AWS does not publish SA goals. What is public: SAs file product feature requests and AWS says most of its roadmap comes from customer feedback. What SA write-ups and promotion guides describe (inference, not official):
Technical wins and adoption. Workloads that land on AWS and grow because of your design work.
Customer outcomes. Features shipped, launches that happened, problems solved.
Influence on AWS. PFRs filed and championed, feedback that changed a service.
Scaling yourself. Reusable content: blog posts, reference architectures, workshops, talks.
The programs and mechanisms you will use
You do not need to recite these. You need to know which one fits which moment. Checked against AWS pages in September 2026. Funding amounts and names change often.
Program
What it is, in one line
Reach for it when
Working Backwards (PR/FAQ)
Amazon's method: write the launch press release and FAQ before building.
A feature idea is vague or everyone wants AI "somewhere."
Experience-Based Acceleration (EBA)
A hands-on sprint: 3 to 6 weeks of prep, ending in a 1 to 3 day build "party" with AWS and partner experts in the room.
The team knows what to build but is stuck or slow on how.
Generative AI Innovation Center
AWS scientists and strategists who take GenAI ideas to production. Launched 2023 with $100M, investment doubled in 2025.
High-value, science-heavy use case with an executive sponsor.
PACE
AWS Prototyping and AI Customer Engineering. Builds GenAI and agentic prototypes with customers in days to weeks.
A strategic feature needs a working prototype to prove feasibility.
Specialist SAs
Domain experts (GenAI, security, data) who back up the account SA.
You hit the edge of your own depth.
Product feature request (PFR)
A structured request from a customer, filed by the SA, to an AWS service team.
A missing feature or region blocks the customer.
ISV Accelerate
Co-sell program: AWS sellers get credit for helping sell the ISV's product. Needs a transactable Marketplace listing and APN membership.
Beacon wants AWS sellers to help it win counties and districts.
AWS Marketplace
Customers buy the ISV's software through their AWS account and contracts, including private offers.
A county wants to buy Beacon using committed AWS spend or an existing contract vehicle.
SaaS Factory
AWS guidance and experts on SaaS architecture: multi-tenancy, tenant isolation, onboarding, SaaS metrics.
Beacon's per-customer deployments are getting expensive and it wants a real multi-tenant design.
Migration Acceleration Program (MAP)
Funding and method for migrations, in three phases: Assess, Mobilize, Migrate and Modernize.
Beacon still runs some products in a data center or on another cloud.
Partner funding (POC, Innovation Sandbox)
POC funding for small customer projects, and sandbox credits for partners building solutions.
Beacon needs credits to prototype without a budget fight.
Well-Architected review
A structured review against AWS's six pillars. The GenAI lens and SaaS lens apply here.
Before any production launch.
Experience-Based Acceleration (EBA)
The problem
Beacon's team has read the docs and watched the videos. Three months later, no feature. Every small blocker (an IAM policy, a VPC endpoint, a GovCloud quirk) costs a week of tickets.
Picture it
A barn raising. The whole community shows up on one day with the materials pre-cut. By sunset the barn stands. It worked because weeks of planning happened first, and every expert was in the same field at the same time.
In plain words
You pick a real outcome, prepare for a few weeks (access, data, backlog, blockers cleared), then put the customer's engineers in a room with AWS and partner experts for one to three days to build it. The customer's people do the work. The experts unblock them on the spot.
The real term
EBA. AWS runs them for migration, modernization, GenAI, and FinOps. Prep runs 3 to 6 weeks and ends in the "EBA party." Output is working software plus a team that knows how to keep going.
At Beacon
A GenAI EBA for Dispatch Assist. Goal at the party: a working pipeline in their GovCloud account that turns a call-taker's notes into a suggested incident code and draft narrative, scored against 200 real (redacted) calls.
Product feature requests (PFRs)
The problem
Beacon needs a specific model or feature in GovCloud and it is not there. If the SA shrugs, Beacon either waits in silence or goes elsewhere.
Picture it
A regular at a small bakery says, "I wish you made rye on Saturdays." One comment is noise. If the counter staff write it down with the name and how often that person buys, and ten regulars say the same, the baker changes the schedule.
In plain words
You write down what the customer needs, why, what it blocks, and how much business depends on it, and you send it to the service team that owns it. Then you follow up, and you tell the customer plainly what happened.
The real term
PFR. SAs file them. AWS says about 90 percent of its roadmap comes from customer feedback. A strong PFR has the customer impact, the workaround cost, the revenue at stake, and a clear ask.
At Beacon
"Beacon (300 public safety agencies, of which 40 require GovCloud [example]) needs model X in GovCloud with the same feature set as commercial. Without it, Dispatch Assist launches commercial-only and 40 agencies wait. Workaround: model Y, which scored 6 points lower on their eval set."
Innovation Center vs PACE vs ProServe
The problem
Three AWS teams all "help build GenAI." Pick the wrong one and you waste the customer's time and burn goodwill with an internal team.
Picture it
Building a house. The architect's research lab (Innovation Center) designs a new kind of roof that has never been built. The model-home crew (PACE) builds one sample room fast so you can walk in it. The contractor (ProServe or a partner) builds the whole house on a paid contract.
In plain words
Science-heavy and high value: Innovation Center. Prove it works with a real prototype in weeks: PACE. Build it at full scale for pay: ProServe or a partner. Your SA work runs through all three.
The real term
AWS Generative AI Innovation Center, AWS Prototyping and AI Customer Engineering (PACE), AWS Professional Services, AWS Partners. The Innovation Center says 73 percent of its initiatives move from proof of concept to production.
At Beacon
Digital evidence search across video transcripts may need the Innovation Center (hard retrieval and eval science). Dispatch Assist fits PACE (clear scope, needs a prototype). Migrating the permitting portal to multi-tenant fits a partner under MAP.
Working Backwards
This is the heart of Bill's team. If you can run a working backwards session well, you are doing the job on day one.
Working Backwards and the PR/FAQ
The problem
Beacon's CEO says "we need AI in the product." Engineering builds a chatbot in the corner of the dispatch screen. Dispatchers never click it, because it solves a problem nobody had. Six months gone.
Picture it
Planning a trip by writing the postcard first. "Dear Mom, sat on the beach in Goa, ate fresh fish, no work email for 7 days." Now you know what the trip must deliver, and you can see that the cheap flight with two layovers does not fit.
In plain words
Before building, write the press release you would publish on launch day, in the customer's words. Then write the FAQ: the hard questions customers and executives would ask. If the press release is boring or the FAQ has no good answers, you just saved months.
The real term
Working Backwards, with the PR/FAQ document. The press release is about one page. The FAQ splits into external (customer questions) and internal (feasibility, cost, risk, metrics). Amazon uses it for its own products, and AWS runs it with customers.
At Beacon
Instead of "add AI to dispatch," the PR/FAQ lands on "Dispatch Assist cuts the time a call-taker spends coding an incident, and never sends a unit on its own." That sentence shapes the whole design: assistive, human confirms, measured in seconds saved.
Go deeper (for follow-up questions)
The document is a forcing function for clear thinking. Most PR/FAQs never become products, and that is the point: cheap to kill on paper, expensive to kill in code.
The internal FAQ is where an SA adds the most: "What does it cost per call?", "What happens when the model is down?", "Which models are available in GovCloud?", "How do we measure accuracy before launch?"
The five customer questions
#
Question
Beacon's answer for Dispatch Assist
1
Who is the customer, and what do we know about them?
The 911 call-taker and dispatcher at a mid-size county. High stress, two screens, typing while talking, many are new hires.
2
What is the customer problem or opportunity?
Coding the incident type and writing the narrative takes time during the call, and new call-takers pick wrong codes, which sends the wrong resources.
3
What is the solution and the most important benefit?
Live suggestions for incident code, premise hazards, and a draft narrative. Benefit: faster, more consistent coding with the human in charge.
4
How do we describe the solution and experience?
"A second set of eyes that never gets tired." Suggestions appear in a side panel. One keystroke accepts. Nothing happens without the call-taker.
5
How do we test it and measure success?
Shadow mode on recorded calls first. Measure code agreement with supervisors, seconds to code, acceptance rate, and any case where a suggestion was wrong and accepted.
How you would run a 90-minute session with Beacon's product leaders
Before: send a two-line pre-read ("Bring one dispatcher story that went badly. Bring your top three AI ideas."). Ask for a product lead, an engineering lead, someone who has sat in a 911 center, and someone from compliance.
Minutes
Block
What you do
0 to 10
Frame
Goal of the day: one draft press release for one feature. Rules: talk about customers, not technology. No model names until minute 60.
10 to 25
Who is the customer
Pick one persona. Collect real stories. "Tell me about a call that went wrong." Write quotes on the board.
25 to 40
The problem
List problems from the stories. Vote. Pick the one that is frequent, painful, and measurable.
40 to 55
Solution and benefit
Brainstorm solutions to that one problem. Push for the smallest version that delivers the benefit. Ask "what would make a dispatcher refuse this?"
55 to 75
Write the press release
Draft together: headline, subhead, problem, solution, leader quote, customer quote, how to start. Short sentences, customer words.
75 to 85
Hard questions
Fast list of external and internal FAQs. You add the technical ones: GovCloud, cost per call, failure mode, evaluation.
85 to 90
Next steps
Owner for the PR/FAQ. Date for the architecture session. What data you need (for example, 200 redacted calls for an eval set).
Sample press release: Beacon Dispatch Assist
Fictional, for practice. Read it aloud once. It takes about two minutes.
Beacon launches Dispatch Assist, a second set of eyes for 911 call-takers
Available today to Beacon CAD customers, including agencies that run in AWS GovCloud (US). Call-takers stay in control of every decision.
RIVERTON, OHIO, March 2027. Beacon today launched Dispatch Assist, a feature inside Beacon CAD that helps 911 call-takers code incidents and write call narratives while they are still on the phone.
A call-taker handling a medical call types, listens, and codes at the same time. New hires take months to learn hundreds of incident codes, and a wrong code can send the wrong unit. Supervisors spend hours each week fixing narratives before they go to records.
Dispatch Assist reads the call-taker's notes as they type. It suggests the incident code, shows known hazards at the address from the agency's own premise history, and drafts the narrative. The call-taker accepts with one keystroke or ignores it. Dispatch Assist never dispatches a unit and never contacts a caller. Every suggestion and every decision is logged for review.
"Our customers told us they did not want a robot dispatcher. They wanted help with the paperwork so they can listen to the caller," said Dana Ortiz, CEO of Beacon. "Dispatch Assist does that, and it runs inside the same secure environment their CAD already uses."
In a 90-day pilot at a county 911 center, call-takers accepted the suggested code on most calls and supervisors reported fewer narrative corrections. [Pilot numbers to be filled in from the real pilot.]
"On my third week, I had a caller describing chest pain and a fall. Dispatch Assist showed the right code and flagged a dog at the address. I still made the call, but I was not guessing," said a call-taker at the pilot center.
Agency administrators can turn on Dispatch Assist from the Beacon admin console, choose which call types it covers, and review a weekly accuracy report. Agency data stays in the agency's Beacon environment and is not used to train any model.
Three external FAQs a 911 director would ask
Can it dispatch on its own? No. It only suggests. The call-taker makes every decision and every suggestion is logged.
What happens if it is down? The CAD works exactly as before. Dispatch Assist is a side panel, not in the critical path.
Is our call data used to train AI? No. Data stays in your environment. The model provider does not see it and it is not used for training.
Four internal FAQs Beacon's CFO and CTO would ask
What does it cost per call? Estimate tokens per call times model price, then test on real calls. Route simple codes to a small model. Target a cost per call below what a supervisor spends fixing narratives.
How do we know it is accurate before launch? An eval set of redacted real calls with supervisor-approved codes. Run shadow mode for weeks before any agency sees it.
Can we run it in GovCloud? Check the current Bedrock model list for GovCloud and pick from it, with a model abstraction so we can switch later.
What is the worst failure? A wrong code accepted without thought. Mitigations: show confidence, never auto-fill high-risk codes, audit accepted suggestions weekly.
The presentation round
What to expect
Reports from candidates and prep sites agree on the shape, and disagree on the details. Plan for this and confirm with the recruiter:
Length: 30 to 60 minutes in the slot, of which your talk is 20 to 30 minutes and the rest is questions. Several candidates describe a 1-hour presentation block inside a 5-hour loop.
Topic: usually your choice: "present a project you delivered: the challenge, your architecture decisions, the outcome." Some teams give a scenario instead, or a written work sample. Some send a slide template.
Audience: a mixed panel of SAs, an SA manager, sometimes an account manager. Treat them as a customer: part technical, part business.
Interruptions: expected and deliberate. They push past your scope to see how you handle challenge without getting defensive.
What they score: clear story (problem, constraints, decisions, outcome), technical depth that survives questions, and the ability to adjust to technical and non-technical listeners. Leadership Principles come up here too.
How to pick a topic
Test
Why it matters
You built it and can defend every decision
They will drill. Anything you did not decide, you cannot defend.
It is GenAI with a real design choice
The role is GenAI, agentic, and data. A pure web app misses.
It has a number you measured
SAs are asked "how do you know it works?" all day.
It maps to a regulated customer
Their customers are 911 centers and school districts. Trust is the hard part.
You can show it running
A live demo in minute 2 buys attention for minute 20.
It includes a failure you owned
Shows judgment and Earn Trust. Every panel loves a real miss.
Recommendation: Ask the Declaration
Of your projects, askthedeclaration.com is the best fit for this panel, for five reasons:
It is a public sector problem. Citizens asking questions about a founding legal document. A wrong paraphrase of the law is worse than no answer. That is the exact fear of a 911 director or a permitting office.
It has a measured eval. A gold question set, retrieval accuracy, and a measured gap between real answers and off-topic questions that sets an abstain threshold. The v2 upgrade moved recall from 93.3 to 97.8 percent with hybrid search and a reranker.
It has real design trade-offs. Retrieval-only vs generative, on-device vs hosted, small vs large embedding model (you measured bge-base and rejected it), opt-in generation because groundedness was only about 61 percent on the fallback model.
It is live and public. You can open it on screen. It is a public repo. The launch post drew 55K+ views.
It translates to AWS for Beacon. The same design becomes Bedrock Knowledge Bases with hybrid search, a reranker, and a grounding check, for Beacon's permitting portal. That slide proves you can take a lesson to a customer.
Why not the others as the lead: the Kendra search at Prudential is your strongest production story, but you cannot show it, so it goes inside the talk as proof of scale (slide 9). PROBE is excellent but abstract for a mixed panel, so it is the backup. The Bedrock CRM prototype is small; use it in objection handling and behavioral answers.
Title and one-line promise
"Answers you can cite: building a GenAI feature that knows when to say I don't know."
"In 25 minutes I will show you a live app I built, how I measured it, what broke, and how I would build the same thing on AWS for a GovTech customer whose answers have to hold up in front of a city council."
Slide by slide (11 slides, about 25 minutes)
Slide 1. Title and why you should care (1 minute)
Title, your name, one line: "I lead engineering teams in regulated financial services and I build GenAI apps on weekends."
The promise sentence above.
"Every one of your GovTech customers wants AI. Most of them are afraid of one thing: a confident wrong answer with their logo on it. This talk is about that fear, and a design that answers it."
Slide 2. The customer and the problem (2 minutes, plus live demo)
Customer: a citizen or student who wants to know what the Declaration or the Constitution says, without reading all of it.
Problem: chatbots paraphrase and invent. For a legal text, a paraphrase is a liability.
Live demo: ask one good question, then one off-topic question ("What is the best pizza in New York?") and show it declines.
"Watch the second question. It does not guess. That refusal is the most important feature in this app, and it is also the one I had to measure hardest."
Slide 3. Working backwards: the tenets (2 minutes)
The press release sentence: "Ask a question in plain English and get the exact passage, with its source, in under a second."
Four tenets, in order of priority: cite the source; say I don't know rather than guess; cost nothing per question; work offline.
"I wrote these tenets before any code. When two of them fight, the higher one wins. That is how I decided every trade-off you will see next."
Slide 4. Architecture v1 (3 minutes) Whiteboard here
Draw it: passages chunked by structure, embedded once at build time, shipped with the app; query embedded in the browser with MiniLM via Transformers.js; cosine similarity; return top matches.
No generation step. Service worker for offline.
Trade-off table: retrieval-only (no hallucination, no per-query cost, cannot combine articles) vs generative (can synthesize, can invent, costs per call).
Slide 5. How I measured it (3 minutes)
Gold set: questions paired with the correct passage, checked against the source text first.
Metrics: top-1 and top-3 accuracy, answer retrieval, and the honesty gap: real answers scored 0.55 to 0.77 similarity, off-topic capped at 0.35. The abstain threshold sits in that gap.
The eval runs the same model the browser runs, so you measure the shipped system.
"The number I care about most is not accuracy. It is the gap between 0.35 and 0.55. That gap is what lets the app say I don't know with confidence."
Slide 6. What was not good enough, and v2 (3 minutes)
Pure vector search missed exact-word questions and some multi-part questions.
Change: hybrid retrieval (BM25 keyword plus vectors) and a cross-encoder reranker (ms-marco MiniLM).
Result: recall 93.3 to 97.8 percent; multi-hop slice 100 percent.
Rejected: a bigger embedding model (bge-base). You measured it and it did not pay for its size in the browser.
Slide 7. Adding generation, carefully (2 minutes)
Opt-in generation in the browser with WebGPU (Llama 3.2 1B via WebLLM), feature-detected, with a fallback.
Groundedness eval: about 61 percent on the fallback model. Not good enough to be the default.
Decision: generation stays opt-in and every answer shows the cited passages.
"This is the slide I would show a customer's legal team. We measured the generative answer, it was grounded about 61 percent of the time on the small model, so we did not ship it as the default. The data made the decision, not the demo."
Slide 8. The same design on AWS, for Beacon (4 minutes) Whiteboard here
Switch customers: Beacon's permitting portal. A resident asks "Do I need a permit for a 7-foot fence?" The answer must cite the county's ordinance, or say "call the permit office."
The browser design translated to managed AWS pieces: hybrid retrieval, rerank, generation, grounding check, and an eval gate.
Ordinances into S3, then a Bedrock Knowledge Base with a vector store that supports hybrid search (OpenSearch Serverless is the usual pick).
Rerank the top passages with a Bedrock rerank model.
Generate through the Converse API so the model can change without code changes.
Guardrails contextual grounding check blocks answers not supported by the passages. Below threshold: decline and route to a human.
The eval set runs in CI. A change that drops recall or grounding does not ship.
Per-county data separation: one knowledge base or metadata filter per tenant.
Slide 9. I have done this in production (2 minutes)
At Prudential you led an Amazon Kendra semantic search integration from POC to production for 35,000+ advisors, replacing keyword search and cutting 20+ steps from advisor workflows.
Two lessons that carry to Beacon: permissions belong in the retrieval layer (a user only retrieves what they can see), and adoption needs time with the real users, not only good relevance.
"The weekend app taught me how to measure. The Prudential rollout taught me that a search nobody adopts is a failed project, so I spent real time with operations people until it fit their day."
Slide 10. How I would run this with a customer (2 minutes)
Week 1: working backwards session. One feature, one press release.
Week 2: build the eval set with the customer's experts (for permitting, 100 real resident questions with the right ordinance).
Weeks 3 to 4: prototype against the eval set. PACE or an EBA if the feature is strategic.
Week 5: Well-Architected review with the GenAI lens, security review, cost per question estimate.
Then the customer's team builds for production, with you on call for design decisions.
Slide 11. Close (1 minute)
Three takeaways: measure before you ship; the ability to decline is a feature; design so the model can change.
"I would love your questions, and especially where you would push back."
Timing plan
Slides
Minutes
If you are running long, cut
1 to 3
5
Tenets become one sentence
4 to 7
11
Slide 7 becomes one sentence on slide 6
8
4
Never cut. This is the slide for this job.
9 to 11
5
Slide 10 becomes a verbal list
Ten likely interruptions, with answers
1. Why not use a large LLM with a long context window and skip retrieval?
For one document I could, but I would pay per question, I would lose the exact citation, and I would have no way to say I don't know with a measured threshold. For Beacon's permitting case it is worse: every county has thousands of pages of ordinances, they change, and each county must only see its own. Retrieval gives me citations, tenant separation, and freshness. Long context is a good tool for a single long document a user uploads.
2. How did you pick the 0.35 and 0.55 numbers? Isn't that overfit to your gold set?
I measured them. Real answers scored 0.55 to 0.77 and off-topic questions topped out at 0.35, so the threshold sits in that gap. It is fit to my data, which is the point, and it would need re-measuring for any new corpus or model. For Beacon I would build the eval set per customer type and re-check the threshold whenever the embedding model changes. I would also add a small set of adversarial questions that look on-topic but are not.
3. Your eval set is small and you wrote it. How do you trust it?
I trust it for regression, not for a marketing claim. It catches a change that makes things worse. I verified every gold answer against the source text first, and I documented the misses instead of rewording questions to pass. For a customer I would have their domain experts write and label the questions, and grow the set from real user questions and from every correction in production.
4. How does this change in GovCloud?
Three things. The model list is shorter, so I pick from what is authorized there and keep the model behind the Converse API so I can switch. Some features arrive later, so I check each piece (knowledge bases, rerank, guardrails) before I promise the design. And the compliance scope is stricter, so logs, encryption keys, and network paths (PrivateLink, no public endpoints) are part of the design from day one.
5. What does this cost per question on AWS?
I would not guess on stage. I would estimate it: tokens in (question plus retrieved passages) and tokens out, times the model price, plus retrieval and rerank costs, and then measure on 100 real questions. The levers are fewer and shorter passages after reranking, a smaller model for simple questions, and caching frequent questions. For permitting, the comparison is the cost of a phone call to the permit office.
6. What about prompt injection? An ordinance PDF could contain instructions.
Treat every retrieved passage and every user message as data, not instructions. Keep the system prompt narrow, give the model no tools that can change anything, and put permissions in the data layer so an injected instruction cannot reach another tenant's data. Guardrails can filter prompt attacks on input. In this design the model can only answer or decline, which limits the damage.
7. Why a reranker? Isn't that extra latency?
Yes, some. Vector search is fast and good at recall but ranks loosely. The cross-encoder reads the question and passage together and ranks precisely. I retrieve a wider set cheaply, rerank the top few, and send fewer passages to the model, which often saves more tokens than the rerank costs. On my set it moved recall from 93.3 to 97.8 percent. I would measure the latency budget for Beacon before choosing.
8. How would you make this agentic?
Only where it earns it. A resident question that needs a lookup (parcel zoning from the county GIS) is a good first tool. I would give the agent read-only tools, a step budget, and a log of every tool call. Anything that changes state, like filing a permit application, needs the resident to confirm. Bedrock Agents or AgentCore can host that, and MCP is a clean way to expose Beacon's APIs as tools.
9. What would you do differently?
Build the eval set with outside users earlier. My gold set reflects how I ask questions. Real users phrase things differently, which is where the cross-lingual and multi-part misses came from. For a customer, I would start collecting real questions in shadow mode before tuning anything.
10. How would you convince a skeptical county attorney to allow this?
Show the mechanism, not the demo. Every answer cites the ordinance section. Below a measured threshold it declines and gives the permit office number. Nothing is stored beyond what the county approves, and nothing is used for training. Then offer a shadow period where staff review answers before residents see them, with a weekly report. Attorneys respond to controls they can inspect.
Backup topic: PROBE, measuring AI behavior
Use this if the panel wants something more enterprise, or if the Declaration demo cannot run. PROBE is live at framework.swapniltamse.com. Title: "Turning 'be fair' into a number you can defend."
Slide
Content
1. Problem
Policies say "be fair, be grounded, resist injection." Nobody can tell you if the chatbot does it. Beacon's school districts will ask exactly this about Beacon Learn.
2. Method
Five steps from principle to probes to rubric to judge to evidence. PROBE is the instrument for the "measure" step of a governance process.
3. Architecture (whiteboard)
Target model and judge behind interfaces. FastAPI service. Claude as judge. Four principles: gender fairness, groundedness (threshold 0.85), prompt injection, refusal.
4. Testing a non-deterministic system
Fakes for target and judge so the core test suite makes zero live calls. 32 tests passing. Live judge behind an integration marker.
5. Demo
The sample flawed bot scoring 41/100 on fairness, with the evidence shown.
6. Security
"Connect your API" mode runs probes server-side with an SSRF guard (https only, private addresses blocked).
7. On AWS for Beacon
Bedrock model evaluation or a Step Functions pipeline running the probe set on every model or prompt change, results to S3 and a dashboard for district IT.
8. Limits and close
The classical-ML evaluator is designed, not shipped. Judges need their own reliability checks.
Customer conversations
The panel may run a mock customer meeting, or ask "how would you start with a new ISV?" Your edge: you have sat on the customer side for 15 years. Use it. You know what it feels like when a vendor talks at you.
Discovery questions: CTO vs engineering lead
Topic
Ask the CTO
Ask the engineering lead
Goals
What does the board expect from AI in the next two quarters? What would make this year a win?
What is on your roadmap that you are most worried about delivering?
Customers
Which customers are asking for AI, and what exactly are they asking for? Which ones are scared of it?
What do support tickets say users struggle with most?
Competition
Who are you losing deals to, and why?
Have you tried anything with GenAI yet? What happened?
Platform
Where is the product going architecturally in three years?
Walk me through the architecture of the product you would add AI to. Single-tenant or multi-tenant?
Data
Who owns the data, you or the agency? What do contracts say?
Where does the data live, how clean is it, and can you get a labeled sample?
Compliance
Which certifications are blocking deals (CJIS, FedRAMP, StateRAMP/GovRAMP, state privacy laws)?
Which customers need GovCloud? How do you deploy there today?
Team
Do you plan to hire for AI, or grow the team you have?
Who on the team would own evaluation and prompts?
Money
How do you plan to price AI features: bundled, add-on, per use?
What is your current AWS spend profile, and what are you told to cut?
Decision
Who else needs to say yes? Legal, security, a customer advisory board?
What would make you say this design will not work?
How to run a first meeting
Before: read their website, product docs, public case studies, job posts (they reveal the stack), and any AWS account notes. Talk to the account manager for 15 minutes.
Open (5 min): introductions, confirm the time, and agree the goal. "My goal today is to understand where you are going, and leave with one next step we both think is worth doing."
Listen (25 min): discovery questions. Take notes in their words. Ask one follow-up for every answer.
Reflect back (5 min): "What I heard is... did I get that right?" This is where trust starts.
Offer (10 min): one or two ideas, each with a trade-off. No slides unless they ask. A quick whiteboard is fine.
Close (5 min): one concrete next step with an owner and a date. For example, a working backwards session in two weeks.
After (same day): a short email with what you heard, the next step, and anything you promised. Then do what you promised.
Qualifying a GenAI use case
The problem
Beacon brings you twelve AI ideas. They can staff two. If they pick the flashy one with no data and a scary risk profile, they burn both quarters and the CEO decides "AI does not work for us."
Picture it
A doctor in an emergency room doing triage. Not every patient gets seen first. The doctor checks a few vital signs fast and sorts. Same with ideas: four quick checks, then sort.
In plain words
For each idea ask four things. Is it worth a lot to a real user? Can current models do it well enough? Is the data there and allowed to be used? What happens when it is wrong? Start with ideas that score well on all four, not the most exciting one.
The real term
Value, feasibility, data readiness, risk. Some teams add a fifth: time to value. The output is a prioritized backlog and a first use case with a success metric.
At Beacon
A narrative draft for dispatch scores high on value and data, medium on risk (a human reviews). An AI that recommends which unit to send scores high on value but very high on risk. Start with the first.
Lens
Questions to ask
Red flag
Value
Who uses it, how often, what does it save or earn? Would a customer pay more or renew because of it?
"It would be cool." No named user or metric.
Feasibility
Can a current model do this on 20 real examples today? Is latency acceptable? Is the model available in the regions needed?
Needs near-perfect accuracy with no human review.
Data readiness
Does the data exist, is it clean enough, do contracts allow this use, can we build an eval set?
Data belongs to the agency and the contract is silent on AI use.
Risk
What happens when it is wrong? Who is harmed? Is a human in the loop? Could it create bias against a group?
Wrong output affects a person's liberty, safety, or benefits with no human review.
Beacon's portfolio, sorted:
Idea
Value
Feasible
Data
Risk
Call
Dispatch narrative draft
High
High
High
Medium
Start here
Permitting Q&A with citations
High
High
High (public ordinances)
Low to medium
Start here
Digital evidence video search
High
Medium
Medium
High (CJIS, chain of custody)
Prototype with Innovation Center
Benefits eligibility auto-decision
High
Medium
Medium
Very high
Assist caseworkers only, never decide
Beacon Learn AI tutor for K12
High
High
Medium
High (minors, FERPA, COPPA)
Teacher-facing first, then students with guardrails
Objection handling
The pattern for every objection has four steps. Practice it until it is automatic.
Acknowledge. Say the concern back. It is almost always reasonable.
Ask. One question to find the real concern under it.
Answer with a mechanism. A specific control, number, or design, not a promise.
Offer a next step. Small, concrete, low risk for them.
"Why not Azure OpenAI?"
"Fair question, and a lot of teams start there. Can I ask what draws you to it? Is it a specific model, or something your team already knows?"
"Here is what I would weigh for Beacon. Your product and your customers' data already run on AWS, including in GovCloud. Keeping the model call in the same account means the same IAM, the same KMS keys, the same CloudTrail logs, PrivateLink with no public internet, and one compliance boundary to explain to a county CIO. Bedrock also gives you a choice of models, including Anthropic, Amazon, Meta, and OpenAI models, behind one API, so you are not betting the product on one provider."
"My suggestion: take 50 real examples, run them on two or three models, and let your eval set decide. I will help set it up."
"Will AWS train on our data?"
"No. Bedrock does not use your prompts or completions to train any model, and it does not share them with model providers. The providers do not have access to your prompts or to the Bedrock logs. Your data is encrypted in transit and at rest, and you can keep traffic private with VPC endpoints."
"Is this coming from your legal team or from one of your agencies? I can send the Bedrock data protection documentation, and I am happy to join a call with them to go through it line by line."
"We tried a POC and it hallucinated."
"That is common, and it is fixable in most cases. Can you show me what it did? What question, what data it had, and what it said?"
"Usually one of three things went wrong. Retrieval pulled the wrong passage, so the model made something up. Nobody measured it, so there was no way to tell good from bad. Or the model was asked to answer when it should have declined."
"In my own project I built a gold set of questions and measured the gap between real answers and off-topic ones, then set a threshold so the app says I don't know below it. On AWS the same controls exist: better retrieval with hybrid search and reranking, citations in every answer, a Guardrails grounding check, and an eval set in CI. Let us take 50 of your failed questions and see which of the three it was."
"It's too expensive."
"Compared to what? Can you tell me what you measured? Cost per request, or the monthly bill from the POC?"
"The number that matters is cost per task compared to what the task costs today. If a draft narrative saves a supervisor five minutes, the model can cost a lot less than five minutes of their time and still be a clear win. Then we bring the cost down: route simple cases to a smaller model, shorten prompts, send fewer passages after reranking, cache repeated questions, and use batch inference for anything that does not need to be real time."
"Let us model the cost per call on 100 real calls. I will bring a spreadsheet, you bring the call volume."
"Our police customers will never allow AI."
"Some will not, and they should have that choice. Can I ask which part worries them? Evidence integrity, decisions about people, or data leaving their control?"
"We design for the cautious customer. The feature is off by default, agency admins turn it on per call type. It assists and never decides. Every suggestion and every acceptance is logged. The data stays in the agency's environment and meets their CJIS requirements. Start with work nobody fears: narrative drafts, records summaries, report formatting."
"Would one chief or 911 director be willing to be a design partner? One respected early customer changes the conversation for the other 299."
"We don't have ML engineers."
"You may not need them for this. Calling a model on Bedrock is an API call. Your application engineers can build it. What you do need is one person who owns quality: the eval set, the prompts, and the metrics. That is closer to a senior engineer with a testing mindset than a data scientist."
"To get your team moving: an EBA where your engineers build it with AWS experts in the room, and AI-assisted coding tools like Kiro or Amazon Q Developer to speed up the build. If a use case needs real science, like fine-tuning, that is when we bring in the Innovation Center."
"Vendor lock-in."
"Reasonable worry. Some lock-in is real with any platform, so let us decide where you accept it and where you keep a seam."
"Keep your assets portable: your data in S3 in open formats, your prompts and eval set in your own repo, tracing through OpenTelemetry. Put the model behind one interface, like the Bedrock Converse API or your own thin wrapper, so switching a model is a config change. Use open standards like MCP for tools. The thing worth the most that you build is the eval set, and that is yours no matter where it runs."
"Which part would hurt most to move? Let us design that part first."
"Our customers are in GovCloud and the model we want isn't there."
"That happens, and I will not pretend otherwise. Which model, and what does it do for you that others do not?"
"Three moves. First, check today's list: the Bedrock GovCloud catalog has grown a lot, and it now includes Claude, Llama, OpenAI, Mistral, and Amazon models, among others. Second, run your eval set on the best available one. Often the gap is smaller than expected, and the design keeps the model swappable. Third, I will file a feature request with your numbers: how many agencies, what revenue, what the workaround costs. I cannot promise a date, but I will champion it and tell you what I hear."
"Whether some non-CJIS workloads could run in a commercial region is your and your customers' compliance decision, not mine. I can lay out the options for your compliance team."
"Legal says no."
"Can I meet legal? In my experience 'no' usually means 'not with what we know today.' What exactly are they worried about: data use, liability for wrong answers, IP, or customer contracts?"
"For each concern there is a document or a design answer. Data use: the Bedrock data protection terms. Wrong answers: human review, citations, and logs. Contracts: we can help draft the AI section of your customer agreement and a plain-language AI disclosure for agencies. Many legal teams say yes to an internal pilot first, with no customer data. That gives them evidence."
"Can you just build it for us?"
"I will build with you, and I will build pieces to prove a point: a reference design, a prototype, a spike. But I should not be the owner of Dispatch Assist. It is your product. When a 911 center calls at 2 AM, your engineers need to know every line."
"If you need build capacity, there are good options: PACE for a fast prototype if it qualifies, an EBA where your team builds with experts in the room, or AWS Professional Services or a public sector partner under a contract. I will stay close either way."
"Two quarters is not realistic for this."
"It might not be for everything. What does the CEO need to see in two quarters: a feature in every customer's hands, or one feature live with a design partner and proof it works?"
"A realistic plan: quarter one is working backwards, the eval set, and a prototype. Quarter two is one feature in production with two or three design-partner agencies, plus a roadmap for the rest. That is a real launch and it protects the brand."
"School districts won't let an AI talk to students."
"Many will not, at first. Is the worry student data, safety of what the AI says, or teachers losing control?"
"Start with teachers, not students: lesson plan drafts, feedback suggestions, reading-level adjustments. The teacher reviews everything. Student data stays in the district's tenant, is not used for training, and the design follows FERPA and COPPA from day one. When you do add student-facing features, add content filters tuned for minors, topic limits, and a teacher dashboard that shows every conversation."
Influence and hard moments
These map to Leadership Principles like Earn Trust, Have Backbone; Disagree and Commit, and Customer Obsession. Keep each answer calm, specific, and short.
Influencing without authority
The problem
You cannot tell Beacon's engineers what to do. You cannot tell an AWS service team what to build. Yet your job depends on both moving.
Picture it
A trail guide who wants the group to take the longer, safer route. Ordering does not work, they are paying customers. So the guide shows them the loose rock on the short route, tells the story of the last group, and lets them decide. Most choose the safe route.
In plain words
You move people with evidence, their own goals, and trust built over time. Show data, not opinion. Tie your ask to what they already want. Make it easy to say yes with a small first step. Keep your promises so the next ask is easier.
The real term
Influence without authority. At Amazon it shows up in Earn Trust, Dive Deep, and Have Backbone. It is the main skill of the SA role, since you manage no one on either side.
At Beacon
Beacon's lead engineer wants to fine-tune a model on all dispatch data. You think RAG plus an eval set is faster and safer. You do not argue. You offer to run both on 50 calls in a week and let the numbers decide.
Telling a customer no
Setup: Beacon's product VP wants the benefits portal to auto-approve or deny applications with AI, to cut backlog.
"I understand the backlog is hurting families and your customers. I want to help with that. I cannot recommend an AI that makes the final decision on someone's benefits. When it is wrong, a family loses support, and you will not be able to explain why to a state auditor."
"Here is what I can help you build this quarter: the AI reads the application and documents, checks for missing items, and prepares a summary with the rules it thinks apply and why. The caseworker decides in two minutes instead of twenty. Same backlog relief, and a human owns every decision."
Disagreeing with an AWS service team
Setup: the service team says the GovCloud feature Beacon needs is "not a priority." You believe it affects many GovTech ISVs.
"I understand you have to rank this against many requests. I want to make sure you have the full picture. Beacon alone covers 40 agencies that require GovCloud [example], and I have heard the same ask from two other GovTech ISVs this quarter. The workaround costs them a lower-accuracy model and a split code path."
"Could we get 30 minutes with the PM and one customer engineer? If after that it still is not a priority, I will tell the customers plainly and help them plan around it."
That is Have Backbone; Disagree and Commit. You push with data, then you commit and help the customer work around it.
Escalating a blocker
Setup: Beacon's quota increase for a Bedrock model in GovCloud has been stuck for two weeks. Their launch is in three.
Do the homework first: case numbers, what was tried, dates, the business impact in one line.
Tell the customer what you are doing and when they will hear from you.
Escalate through the right path with your manager copied: account manager, then the service team or support escalation.
Write it in three lines: what is blocked, what it costs, what you need by when.
"Blocked: Beacon's Bedrock quota increase in GovCloud West, case open 14 days. Impact: launch to 12 agencies on October 30 slips. Ask: a decision by Friday, or a clear no so we can plan the fallback model."
A customer executive who is angry
Setup: Beacon's CEO calls. Dispatch Assist gave a wrong suggestion in a live demo at a sheriffs' conference, in front of a customer. She says "your AI embarrassed us."
Let her finish. Do not defend or explain in the first two minutes.
Acknowledge the impact, not blame: "That was in front of a customer. I understand why you are angry."
Get facts: "Can you tell me exactly what happened? What was typed and what it suggested?"
Commit to a specific next step and time.
Follow up before the time you promised.
"I am sorry that happened in front of a customer. I want to understand it fully. Can you send me the input and what it showed? By tomorrow at noon I will have my team and yours look at it and tell you whether it was retrieval, the prompt, or the model, and what we change."
"For the next demo, I suggest we run it on a fixed set of scenarios we have tested, and show the confidence score and the human accept step on screen. That turns a scary moment into your best trust story: the call-taker is always in charge."
Role-play drills
Practice these out loud, ideally with someone playing the customer. Set a timer. Record yourself once and listen back for filler and for how soon you start talking about AWS services.
Drill 1. First meeting with Beacon's CTO (10 minutes)
Setup: first meeting. The account manager is with you. The CTO has 30 minutes.
The customer says: "Our CEO wants an AI strategy by next month. I don't know where to start. What are other companies doing?"
What good looks like:
You do not list what other companies do for five minutes. One sentence, then turn it back: "Before I tell you what others do, what are your customers asking for?"
You ask at least five discovery questions from the table, including data ownership and GovCloud.
You reflect back what you heard in two sentences.
You suggest a working backwards session on the top two ideas as the next step, with a date.
You mention no more than two AWS services.
Drill 2. The angry CEO after a failed demo (5 minutes)
Setup: the scenario from the hard moments section. The partner playing the CEO should interrupt you twice.
The customer says: "Your AI embarrassed us in front of the Sheriff of the biggest county we have. I was told this was ready."
What good looks like:
You stay silent until she finishes, then acknowledge the impact in one sentence.
You do not say "AI is non-deterministic" or blame the demo team.
You ask for the exact input and output.
You commit to a specific time for the root cause.
You propose one concrete change for the next demo (tested scenarios, visible confidence and human accept step).
Drill 3. The engineer who wants to fine-tune (8 minutes)
Setup: architecture session with Beacon's principal engineer for digital evidence search.
The customer says: "RAG is a hack. We should fine-tune a model on all our evidence transcripts. Then it will know everything."
What good looks like:
You respect the idea and ask what problem fine-tuning would solve for them.
You explain the trade-offs plainly: fine-tuning teaches style and format well, but it does not give citations, it cannot forget a case that gets sealed or expunged, and it mixes agencies' data into one model. Retrieval keeps data per agency, cites sources, and updates the moment a file changes.
You mention chain of custody and CJIS as reasons citations and per-agency separation matter.
You propose a test: RAG on 50 real queries this week, and fine-tuning later only if the eval shows a gap that retrieval cannot close.
Drill 4. The K12 product lead in a hurry (8 minutes)
Setup: Beacon Learn's product lead wants a student-facing AI tutor live before the spring semester.
The customer says: "Every edtech company has an AI tutor now. We need one for students in eight weeks or we lose renewals."
What good looks like:
You take the renewal risk seriously and ask which districts said what.
You raise the risks without lecturing: minors, FERPA and COPPA, what the tutor says about self-harm or off-topic questions, and district approval cycles.
You propose a working backwards session this week, and a phased launch: teacher-facing tools in eight weeks, a student tutor pilot with two districts after that, with guardrails, topic limits, and a teacher view of every conversation.
You name how you would measure it: teacher adoption, flagged conversations per thousand, and district sign-off.