Leadership Principles and your stories

Most of the AWS loop is behavioral. This page teaches how the round is scored, what each of the 16 principles means in plain words, and which of your real stories answers each one.

About 60 minutes to read end to end. Skim the quick version if you have 3. Rehearse the story bank out loud; reading it is not enough.

The quick version

  • Every interviewer in the loop owns two or three principles. They ask for a real past example, then keep asking "why" and "how" until they reach the part only you could know.
  • Answer in STAR shape: one breath of Situation, one of Task, most of your time on what you did, then a Result with a number. Two to three minutes, then stop.
  • Say "I" for your actions. "We" is fine for the team's outcome, but the interviewer can only score what you did.
  • Your strongest SA stories: Kendra search for 35,000+ advisors, the live C++ to C# trading cutover, the 570K-record Cosmos load, the Bedrock prototype you said should not ship yet, and the eval you refused to rig.
  • Lead with a real mistake when asked about failure: the conference app you shipped without analytics. Own it, name the fix, show it changed how you work.
  • Gaps to fill before the loop: a second Have Backbone story, a stronger Bias for Action story, and an external-customer story. See the story map.

How the behavioral round works

Amazon hires against its Leadership Principles. The technical part of an SA loop checks whether you can design on AWS. The behavioral part checks whether you will act like an Amazonian when nobody is watching. Both count, and a weak behavioral round can sink a strong technical one.

The loop, the probing, and the Bar Raiser

The problem

A company that hires on charm and a good resume ends up with people who interview well and fail on the job. Nobody checked what they did when a real decision got hard.

Picture it

A home inspection before a sale. One inspector owns plumbing, another owns wiring, another the roof. Each one opens the wall in their area and looks behind the paint. Then there is one extra inspector sent by the bank, not the seller. That person does not care how nice the kitchen is. They only care whether this house is better than the houses the bank already owns.

In plain words
  • You meet several interviewers in a row. Each has been assigned two or three principles.
  • They ask "Tell me about a time when..." and then dig. "What did you do next? Why that? What was the data? What did your manager say?" This is called peeling the onion.
  • They type notes as you talk. Later each writes up the evidence per principle and gives a hire or no-hire vote.
  • All interviewers meet in a debrief and argue it out. The Bar Raiser runs that debrief.
The real term

Loop: the set of interviews in one day. Bar Raiser: a trained interviewer from outside the hiring team whose job is to make sure every hire is better than half of the people already in that role. They can block a hire even when the hiring manager wants it. Debrief: the meeting where the votes are defended with written evidence.

Scoring is evidence based. The interviewer needs your words, in their notes, showing the principle. A great feeling in the room without quotable specifics scores low.

Your story

You have been on the other side of this. You hired and promoted engineers at Prudential, and in your June 2026 prep you wrote that you ask candidates about a time they disagreed with a decision that shipped anyway. That is a Bar Raiser question. Use that empathy: give the interviewer the exact sentence they want to type.

Go deeper (for follow-up questions)

What the notes look like. An interviewer writes things like: "Candidate personally chose parallel run over big-bang cutover because a bad cutover gives a trader a wrong price. Reconciled outputs until modules matched under live load." That is a strong Deliver Results and Are Right, A Lot data point. "We migrated a platform and it went well" is not.

How deep they peel. Expect three to five follow-ups per story. By the third, they are testing whether you were there. Details only a participant knows (the error code, the name of the table, what the VP said) are what convince them.

"I" versus "we". Say "we" once to credit the team, then switch to "I" for every decision and action. If you keep saying "we", the interviewer will ask "what did you do?" and you lose a minute.

Always a number. Users, seconds, percent, records, days, dollars. If you have no number, give a before and after in concrete terms ("call-center escalations stopped").

Length. Two to three minutes for the first telling. Stop and let them pull. Long answers crowd out the follow-ups that earn the score.

New stories per interviewer. Each interviewer usually wants fresh stories. The notes are shared at the debrief, so reusing one story across two interviewers looks thin. Plan roughly two stories per interviewer, across different jobs.

The answer shape

STAR, and how to spend your two minutes

The problem

Under pressure most people spend ninety seconds on background, thirty on what they did, and forget the result. The interviewer ends with nothing to score.

Picture it

Telling an insurance adjuster about a fender-bender. Where were you. What were you trying to do. What exactly did you do, turn by turn. What happened to the car. The adjuster does not want the history of the road.

In plain words
  1. Situation (about 15 percent): company, stakes, one number that shows scale.
  2. Task (about 10 percent): what you personally owned. One sentence.
  3. Action (about 60 percent): two or three decisions you made, and why each one. Name the option you rejected.
  4. Result (about 15 percent): the number, what changed for the customer, and what you learned.
The real term

STAR: Situation, Task, Action, Result. Amazon's own interview prep pages recommend it. Some coaches add an L for Learning. For failure questions, the learning is required.

Your story

Your Vantage migration story already has a good Action: "I ran old and new in parallel on live data, reconciled continuously, cut each module over only when it matched, and kept rollback one step away." Three decisions, each with a reason. Every story in the bank below follows that pattern.

Go deeper (for follow-up questions)

The follow-ups you will get, and how to answer each:

  • "What would you do differently?" Always have one real answer ready. Not "nothing". Name a concrete change and say you now do it by default.
  • "What did your manager (or the customer) say?" Quote a reaction if you remember it. If you do not, say what happened next, which shows the reaction ("they asked us to run it for the second chapter too").
  • "What was the data?" Give the number and how you measured it. If an estimate, say so: "about 30 percent, estimated from sprint throughput."
  • "What was the hardest part?" Usually a person or a trade-off, not the code.
  • "Why did you choose that over X?" Name X and its cost. This is where Are Right, A Lot is scored.

When you do not have a perfect story:

  • Pick the closest real one and say so in one line: "The closest example I have is from my own build rather than a client, and here is why it fits."
  • Never switch to a hypothetical ("I would..."). If they want a hypothetical they will ask for one.
  • Take five seconds of silence to choose. "Let me pick the best example" is fine and reads as judgment.
  • A smaller story told with full detail beats a big story told vaguely.

"The short version is: [one-line result]. Here is how it happened." Then STAR.

Leading with the result gives the interviewer the headline to write down, and they listen to the rest as proof.

The 16 principles

The ten SA interviewers ask about most come first. The last six still appear, so have at least one story each. Links in the green rung jump to the full story in the story bank.

The ten asked most in SA loops

1. Customer Obsession

The problem

Teams build what is interesting to them, or what a competitor launched, and the customer gets a feature they never asked for while their real pain stays.

Picture it

A good family doctor asks "where does it hurt, and when?" before writing anything. A bad one hears "headache" and hands you the pill they prescribe to everyone.

In plain words

Start from the person who uses the thing and work backwards to the technology. Measure success by what changed for them.

The real term

"Leaders start with the customer and work backwards. They work vigorously to earn and keep customer trust. Although leaders pay attention to competitors, they obsess over customers." Amazon's tool for this is the Working Backwards press release and FAQ.

Your story

Best: Search advisors would use (Kendra). You sat with operations people, tuned to their real queries, cut 20+ steps from their workflow. Backup: Advisor Portal, 10+ seconds to under 5 or Rostering for school districts (EdTech, lands well with this team).

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you went above and beyond for a customer.
  2. Tell me about a time you had to balance what a customer asked for with what they needed.
  3. Tell me about a time you used customer feedback to change a technical design.

They listen for: you talked to real users, you can name their pain in their words, you measured the outcome on their side, and you said no to a request when it would not help them.

Red flags: "customer" means your internal boss; the result is a technical metric nobody outside engineering cares about; you never met a user.

2. Earn Trust

The problem

People hide bad news, spin their mistakes, and take credit for others' work. Soon nobody believes a status report and every decision needs three meetings.

Picture it

The mechanic who calls to say "you do not need the brake job the other shop quoted." You go back to that mechanic for twenty years.

In plain words

Listen, say the true thing plainly, admit your own errors first, and give credit where it belongs.

The real term

"Leaders listen attentively, speak candidly, and treat others respectfully. They are vocally self-critical, even when doing so is awkward or embarrassing." For an SA, trust is the currency with a customer's CTO: they must believe you when you say "do not use this service for that."

Your story

Best: An eval that refuses to lie. You documented three German misses instead of rewording the test to go green. Backup: The instrumentation I forgot (vocally self-critical) or NIST, where you publicly credited a peer's idea inside the section you drafted.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you had to deliver bad news to a customer or stakeholder.
  2. Tell me about a time you were wrong and had to tell people.
  3. How did you build trust with a team or customer that was skeptical of you?

They listen for: you raised a problem early, before it was forced out; you named your own part in it; the relationship was stronger afterwards.

Red flags: blaming another team; a "mistake" that is secretly a strength ("I cared too much"); trust built by agreeing with everything.

3. Dive Deep

The problem

Leaders who only read dashboards miss that the green metric is measuring the wrong thing, and the outage arrives with no warning.

Picture it

A chef who tastes the sauce instead of trusting the recipe card. When the numbers and the taste disagree, they believe the taste and go find out why.

In plain words

Go to the logs, the data, the code, yourself. Be suspicious when a metric and a real-world report disagree. No task is too small for you.

The real term

"Leaders operate at all levels, stay connected to the details, audit frequently, and are skeptical when metrics and anecdote differ. No task is beneath them." For an SA this is the principle behind every debugging and root-cause question.

Your story

Best: 570,000 records and a wall of 429s. Backup: An eval that refuses to lie, where you found your own metric was measuring agreement instead of correctness.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about the hardest technical problem you debugged. Walk me through how you found the root cause.
  2. Tell me about a time a metric looked fine but something was wrong.
  3. Tell me about a time you had to learn the details of a system you did not build.

They listen for: the specific signal you noticed (an error code, a count that did not reconcile), the steps you took to narrow it down, and how you proved the fix worked.

Red flags: "I asked the team to look into it"; no technical specifics; a fix with no verification.

4. Ownership

The problem

Everyone does their own ticket and nobody owns the outcome. Problems sit in the gap between teams because each one says "not my job".

Picture it

A renter and an owner both notice a slow leak under the sink. The renter mentions it to the landlord someday. The owner fixes it this weekend, because the rotten floor will be their problem in five years.

In plain words

Act as if the whole result is yours, including the parts outside your job description and the long-term cost of today's shortcut.

The real term

"Leaders are owners. They think long term and don't sacrifice long-term value for short-term results. They act on behalf of the entire company, beyond just their own team. They never say 'that's not my job.'"

Your story

Best: The instrumentation I forgot. You designed for reuse before anyone asked, owned the miss, and built a runbook so volunteers are not dependent on you. Backup: Vantage live cutover, where you were Acting CIO and still the engineer accountable for the code.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you took on something outside your area because nobody else would.
  2. Tell me about a time you made a decision that cost something short term for a long-term gain.
  3. Tell me about a project that failed. What was your part in it?

They listen for: you stepped in without being asked, you followed through after launch, and you owned your part of a failure without spreading it around.

Red flags: waiting for permission; "the other team dropped the ball"; ownership that ends at go-live.

5. Deliver Results

The problem

Teams stay busy, run great meetings, and still miss the date, or ship on time with quality so poor it has to be redone.

Picture it

A caterer at a wedding. The oven breaks an hour before dinner. Guests never find out, because the caterer borrowed the venue's grill and changed the menu order.

In plain words

Focus on the few inputs that decide the outcome, hit the date at the right quality, and when a setback hits, adapt instead of explaining.

The real term

"Leaders focus on the key inputs for their business and deliver them with the right quality and in a timely fashion. Despite setbacks, they rise to the occasion and never settle."

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you delivered under a tight deadline with limited resources.
  2. Tell me about a time you hit a major obstacle close to launch.
  3. Tell me about a goal you missed. What happened?

They listen for: a clear goal, the obstacle, your trade-off (what you cut, what you protected), and the measured outcome.

Red flags: result is "it went well"; you hit the date by quietly dropping quality; no setback in the story at all.

6. Invent and Simplify

The problem

Every new need gets another tool, another team, another step. The system grows until nobody understands it and each change takes months.

Picture it

The person who labels every drawer in a messy kitchen so guests stop asking where the spoons are. Nothing new was bought. The questions just stopped.

In plain words

Find a new way to solve the problem, and make the solution smaller, not bigger. Borrow good ideas from anywhere.

The real term

"Leaders expect and require innovation and invention from their teams and always find ways to simplify. They are externally aware, look for new ideas from everywhere, and are not limited by 'not invented here'."

Your story

Best: Killing the phone call (Confluence rebuild, 70 percent faster retrieval, 90 percent fewer support tickets). Backup: conference app as one typed data file, or the retrieval-only choice in the constitution search.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you simplified a complex process or system.
  2. Tell me about the most innovative thing you have built.
  3. Tell me about a time you used an idea from outside your field.

They listen for: something was removed, not added; a measurable drop in effort or cost; you can name where the idea came from.

Red flags: "innovation" that is adopting a trendy tool; the new thing is more complex than the old one.

7. Have Backbone; Disagree and Commit

The problem

People see the flaw in the plan but stay quiet to keep the peace. Or they lose the argument and then drag their feet for months.

Picture it

A co-pilot who says clearly, "Captain, I think we are too low." If the captain listens and decides to continue, the co-pilot flies the approach with full effort, not half.

In plain words

Say you disagree, with data, to the person who decides, even when it is uncomfortable. Once the decision is made, commit to it completely.

The real term

"Leaders are obligated to respectfully challenge decisions when they disagree, even when doing so is uncomfortable or exhausting... They do not compromise for the sake of social cohesion. Once a decision is determined, they commit wholly."

Your story

Best: Saying "not yet" to a demo the VP liked [confirm the VP's reaction]. This is your thinnest principle. See the story map for a second candidate to build.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you disagreed with your manager or a customer. What did you do?
  2. Tell me about a time you committed to a decision you disagreed with.
  3. Tell me about a time you pushed back on a senior leader.

They listen for: you raised it with data, to the right person, respectfully; and either you changed the decision, or you lost and still delivered it fully. Both endings score.

Red flags: you always win (sounds invented); you went around the decision-maker; you "committed" but kept complaining; the disagreement was trivial.

8. Learn and Be Curious

The problem

Senior people stop learning, keep using the patterns from ten years ago, and advise customers with stale answers.

Picture it

A cab driver who has driven the city for twenty years and still checks the traffic app, because the bridge closed yesterday.

In plain words

Keep learning on purpose, especially outside your comfort zone, and turn what you learn into something you ship.

The real term

"Leaders are never done learning and always seek to improve themselves. They are curious about new possibilities and act to explore them."

Your story

Best: The instrumentation I forgot (you changed how you work because of it). Backup: NIST, where you learned a federal standards process and brought trading-floor controls into it, or your self-built AI apps (on-device RAG, Claude with MCP, multi-agent recipes) as proof you rebuilt hands-on skill.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you had to learn a new technology quickly to help a customer.
  2. What is something you learned recently, and how did you use it?
  3. Tell me about a time curiosity led you to a better solution.

They listen for: a specific method for learning, a short time frame, and a real output (a shipped app, a changed design).

Red flags: a list of courses with no application; learning only when forced.

9. Bias for Action

The problem

Teams study a reversible decision for six weeks. The customer waits, the competitor ships.

Picture it

Choosing a restaurant for tonight versus buying a house. You can pick dinner in two minutes, because if it is bad you eat somewhere else tomorrow. You should not buy a house that way.

In plain words

Sort decisions into easy-to-undo and hard-to-undo. Move fast on the first kind. Take calculated risks.

The real term

"Speed matters in business. Many decisions and actions are reversible and do not need extensive study. We value calculated risk taking." Amazon calls these two-way doors (reversible) and one-way doors (not).

Your story

Best: 570,000 records and a wall of 429s (you changed approach instead of waiting for more throughput). This is a gap: build a sharper one before the loop. See the story map.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you made a decision with incomplete information.
  2. Tell me about a time you took a calculated risk.
  3. Tell me about a time you acted without waiting for approval.

They listen for: you named it as reversible, you limited the blast radius, you decided in hours or days, and you checked the result.

Red flags: reckless (a one-way door taken fast); or "I gathered more data" as the whole story.

10. Are Right, A Lot

The problem

Confident leaders who never test their own view make the same expensive mistake twice.

Picture it

A good weather forecaster. Not right every day, but right often, because they check several models and adjust when the radar disagrees with their gut.

In plain words

Good judgment comes from asking people who see it differently and trying to prove yourself wrong before you commit.

The real term

"Leaders are right a lot. They have strong judgment and good instincts. They seek diverse perspectives and work to disconfirm their beliefs."

Your story

Best: Vantage, where you kept C++ on the latency hot paths instead of migrating everything. Backup: Kendra (managed service over building retrieval from scratch) or retrieval-only over generative.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a difficult technical decision where you had to choose between two good options.
  2. Tell me about a time you were wrong. How did you find out?
  3. Tell me about a time you changed your mind after hearing another view.

They listen for: you named the options, the trade-off, the data you used, and whose view you sought out. Being wrong once and catching it scores well.

Red flags: "I trusted my gut" with no evidence; never having been wrong.

The other six

11. Hire and Develop the Best

The problem

Managers hire whoever is available and then do not grow them. The bar drops with every hire and the best people leave.

Picture it

A youth soccer coach whose players go on to captain other teams. The coach is judged by where the players end up, not by one season.

In plain words

Every hire should raise the average. Coach people, build their case for promotion, and let them go to bigger roles.

The real term

"Leaders raise the performance bar with every hire and promotion. They recognize exceptional talent, and willingly move them throughout the organization. Leaders develop leaders and take seriously their role in coaching others." For an SA, this also covers mentoring customer engineers.

Your story

Best: Growing engineers who ship without me (15+ engineers, 3+ promoted). Backup: NYU Tandon capstone mentoring, or GameDay, where you started building a path to senior for the engineer who kept reframing the problem [confirm].

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about someone you mentored who grew significantly.
  2. Tell me about a hiring decision you got wrong.
  3. How have you raised the bar on a team?

They listen for: a named person (first name only is fine), the gap, what you did over months, and where they are now.

Red flags: mentoring means "I answered their questions"; no concrete outcome.

12. Insist on the Highest Standards

The problem

"Good enough" ships, the defect travels downstream, and the customer finds it.

Picture it

A tailor who unpicks a seam that only they can see is crooked, because it will pull after ten washes.

In plain words

Set a bar others think is too high, keep raising it, and fix problems so they stay fixed.

The real term

"Leaders have relentlessly high standards, many people may think these standards are unreasonably high. Leaders are continually raising the bar... Leaders ensure that defects do not get sent down the line and that problems are fixed so they stay fixed." (Amazon's text uses a dash where the comma is.)

Your story

Best: Break it on purpose (GameDay win turned into org-wide resilience standards). Backup: the eval gated in CI, or the Vantage reconciliation bar.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you refused to compromise on quality.
  2. Tell me about a time you raised the bar for your team.
  3. Tell me about a time you were not satisfied with the status quo.

They listen for: a mechanism (test, standard, gate) that keeps quality up after you leave.

Red flags: perfectionism that missed the date with no customer reason.

13. Think Big

The problem

Teams solve only the ticket in front of them. Each customer gets a one-off, and nothing compounds.

Picture it

Someone asked to fix one pothole who writes down how the city could fix all of them faster.

In plain words

Set a bold direction and explain it so others want to follow. Look for the version that serves many customers, not one.

The real term

"Thinking small is a self-fulfilling prophecy. Leaders create and communicate a bold direction that inspires results. They think differently and look around corners for ways to serve customers." For this role: turning one ISV's pattern into a reusable reference other GovTech ISVs adopt.

Your story

Best: conference app built as a system, not a site (a second chapter asked for it within five days). Backup: NIST.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you proposed a bold idea.
  2. Tell me about a time you turned a one-off solution into something reusable.
  3. What is the biggest idea you have pushed, and how did you get buy-in?

They listen for: scope wider than your role, and proof others adopted it.

Red flags: big vision with no delivery.

14. Frugality

The problem

Every problem gets more headcount and more budget. Costs grow faster than value, and the customer's AWS bill becomes the reason they leave.

Picture it

A camping cook who makes a great meal with one pot, because that is what fits in the pack.

In plain words

Get more done with less. Constraints push you to better designs. No credit for a bigger budget.

The real term

"Accomplish more with less. Constraints breed resourcefulness, self-sufficiency, and invention. There are no extra points for growing headcount, budget size, or fixed expense." For an SA: cost-aware architecture and right-sizing the model for each GenAI call.

Your story

Best: retrieval-only search with zero per-query inference cost. Backup: static app on a CDN with zero hosting cost for a volunteer group, or the Vantage scrapers that replaced paid market-data subscriptions.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you delivered with fewer resources than you wanted.
  2. Tell me about a time you cut cost without hurting the customer.
  3. Tell me about a time you said no to spending money.

They listen for: a design choice that removed a cost line, with the trade-off you accepted.

Red flags: cutting corners on security or quality to save money.

15. Strive to be Earth's Best Employer

The problem

Teams hit numbers by burning people out. The quiet ones never speak up, and the good ones leave.

Picture it

A host who notices the guest standing alone at the party and introduces them to someone.

In plain words

Make work safer, fairer, and more rewarding. Ask whether your people are growing, empowered, and ready for what is next.

The real term

"Leaders work every day to create a safer, more productive, higher performing, more diverse, and more just work environment. They lead with empathy, have fun at work, and make it easy for others to have fun."

Your story

Best: Making acquired engineers feel their work counted. Backup: retention on your Prudential team, or co-leading the Prudential Asian Pacific-Islander American business resource group.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you made your team's work environment better.
  2. Tell me about a time you supported someone who was struggling.
  3. How have you made your team more inclusive?

They listen for: a specific action for a specific person or group, and evidence it helped (retention, promotions, someone speaking up).

Red flags: generic values talk; team parties as the whole answer.

16. Success and Scale Bring Broad Responsibility

The problem

A system used by millions has side effects nobody designed for: leaked data, unfair decisions, a 911 call routed wrong. At scale, small blind spots hurt many people.

Picture it

A camper who leaves the campsite cleaner than they found it, because a thousand other campers will use it this summer.

In plain words

Think about second-order effects on people outside the transaction: communities, the public, future users. Leave things better than you found them.

The real term

"We are big, we impact the world, and we are far from perfect. We must be humble and thoughtful about even the secondary effects of our actions." In GovTech this is responsible AI: a benefits-eligibility model that wrongly denies a family is a real harm.

Go deeper (questions, what they listen for, red flags)

Typical questions

  1. Tell me about a time you considered the wider impact of a technical decision.
  2. Tell me about a time you raised a risk others had not considered.
  3. How do you think about responsible use of AI for public-sector customers?

They listen for: you named a specific harm, designed a control for it, and did so before being asked.

Red flags: "compliance handles that"; safety as an afterthought.

Your story bank

Fourteen stories, built only from your vault notes. Each is sized for a two to three minute first telling. Anything marked [confirm] is a detail the vault does not settle; fix it before you say it out loud. Rehearse each one standing up, with a timer.

S1. The instrumentation I forgot (failure story)

Fits: Ownership, Learn and Be Curious, Earn Trust, Think Big

Situation My professional association's New Jersey chapter runs a 300-attendee conference every year, coordinated by volunteers from a printed booklet and a PDF agenda. I offered to build the day-of app. No budget, no engineering team, a date that could not move, and a venue where the wifi is never reliable.

Task Ship something attendees would use on their phones on the day, and make sure a volunteer group was not left dependent on me afterwards.

Action I made one early call: build a reusable system, not a one-off site. The whole conference lives in one typed data file, so another chapter is a data swap, not a rebuild. I exported it as static HTML on a CDN, so there is no server to patch and nothing to fall over at the 9 AM registration spike. I wrote the service worker myself: pages network-first so the agenda stays fresh, assets cache-first so it opens when the wifi dies. I took it through the national brand review, including an accessibility fix when their orange failed contrast on white.

Result It ran the conference. National marketing approved it and added the DNS record themselves. Five days later a second chapter asked in writing whether we could build the same for them. My mistake: I shipped with the analytics token empty and the feedback endpoint undeployed. Two config values. So I have no usage data and no satisfaction scores from a 300-person event, and the strongest argument, that it worked, is the one I cannot prove. I now treat measurement as a launch deliverable with its own checklist line, set before go-live.

Follow-up they will ask: "How would you fix it now?" A cookieless analytics beacon and a small serverless endpoint for ratings, both free, both set and tested the week before the event, not the night before.

S2. Moving a trading platform while the desk kept trading (deliver under pressure)

Fits: Deliver Results, Are Right A Lot, Insist on the Highest Standards, Ownership

Situation I spent seven years at Vantage Commodities, a live energy and commodities trading firm, and grew from engineer to Acting CIO. Our core trading systems were in C++. A failure there was immediate revenue loss on the desk, not a ticket for tomorrow.

Task Move the core platform from C++ to C# for maintainability, without ever handing a trader a wrong price or a broken session. I owned the decision and the code outcome.

Action I rejected a big-bang cutover. Instead I ran old and new systems in parallel on the same live feed, with a reconciler comparing outputs continuously. A module cut over only once its output matched under real trading load, and rollback stayed one step away. I also made a call some people questioned: I kept C++ on the latency-critical hot paths, because there milliseconds moved money and C# added latency. I migrated the business logic where C# won on productivity. I explained that choice to the traders and the CTO in their terms, execution speed and risk, not a language debate.

Result The migration completed with no disruption to the desk. The cross-language patterns stayed in use long after. Over the same years I also moved the firm from on-prem colocation to Rackspace and then to AWS.

Follow-up they will ask: "How did you know it was safe to cut over?" Continuous reconciliation on live data. A module was trusted only once old and new matched under real load, lowest-risk modules first. [confirm: team size and how long the parallel run lasted]

S3. Search advisors would use (Kendra to production)

Fits: Customer Obsession, Deliver Results, Are Right A Lot, Earn Trust

Situation At Prudential, advisor operations in Global Investment Operations relied on a legacy keyword search. Advisors phrase questions in natural language, so finding the right answer took many steps. The portal served 35,000+ financial advisors.

Task I led the effort to replace it with AI search and get it into production, not a pilot.

Action I chose Amazon Kendra, a managed service, over building retrieval from scratch, because we needed connectors, access control and relevance tuning on a business deadline. Access control lived at the retrieval layer, so search only returned what each user was entitled to see. The integration ran on Node.js, Lambda, EventBridge and Python. Then I spent real time with the operations people: I sat with their workflow, tuned relevance against their actual queries, and showed them the before and after. I had the AWS compliance documentation and rollback criteria ready before legal and risk asked.

Result It went to production and cut 20+ steps from advisor workflows, measured on Prudential's Ease of Doing Business score. The legacy keyword search was retired.

Follow-up they will ask: "Why Kendra and not your own RAG pipeline?" Managed connectors and access control got us to production fast. I would build my own when I need full control over chunking, embeddings and the generation step, or when managed cost at scale outweighs speed. [confirm: timeline from POC to production]

S4. From 10 seconds to under 5 for 35,000 advisors (Advisor Portal)

Fits: Customer Obsession, Dive Deep, Deliver Results

Situation Prudential's Advisor Portal served 35,000+ financial advisors and, through them, millions of end customers. Search took more than 10 seconds, returned poor results, and advisors escalated to the call center.

Task I led the portal's modernization, with fixing the search experience as the visible goal.

Action We broke the monolithic dependencies into 8 Java Spring Boot microservices, 3 front end and 5 back end, on AWS (EC2, S3, EventBridge, Lambda), and built a REST API layer the portal could call directly. I set test-driven development with JUnit and Mockito as the team standard for these services. [confirm: the specific root cause you found for the 10-second searches, for example a query pattern, missing index, or synchronous call chain. The interviewer will ask.]

Result Search response dropped from 10+ seconds to under 5, and the call-center escalations about search stopped.

Follow-up they will ask: "How did you find where the time was going?" Have one concrete tracing step ready (New Relic was in the stack) and the single biggest fix. [confirm]

S5. An eval that refuses to lie

Fits: Dive Deep, Insist on the Highest Standards, Earn Trust, Frugality

Situation I built and shipped a browser-based search over Liechtenstein's constitution, German and English, 122 articles. It cites the exact article and never generates law. "It looks right on my queries" is not evidence.

Task Prove three claims with numbers I could defend: it finds the right article, it works in both languages, and it refuses off-topic questions instead of inventing an answer.

Action I chose retrieval-only on purpose: the right answer is the source passage, so a generation step would add cost and hallucination risk. Then I built an eval that loads the same embedding model the browser uses, so I tested the real system. Three calls mattered. I found my first cross-language metric measured whether two queries agreed, not whether either was correct, and rewrote it. I set the pass bar just under the measured baseline so it guards against regressions instead of failing every run. And when three German phrasings retrieved an adjacent article, I documented the misses instead of rewording the test set to go green. A fast deterministic suite runs in about a second on every push, and CI is gated on both.

Result Top-1 88 percent, top-3 100 percent, answer retrieval 100 percent, cross-language 63 percent with every miss written down. Real answers score 0.55 to 0.77 and off-topic queries cap at 0.35, so the tool abstains instead of hallucinating. Zero per-query inference cost, no backend.

Follow-up they will ask: "Why ship at 63 percent?" Because the bar is a regression guard. A 90 percent bar failed every run and would be ignored within a week. The English side is strong; the German misses are documented and on the list.

S6. 570,000 records and a wall of 429s (dive-deep debugging)

Fits: Dive Deep, Bias for Action, Deliver Results

Situation On a client engagement in regulated asset management, we had to load 570,000+ records from CSV into Azure Cosmos DB. The first, one-record-at-a-time version kept dying partway through.

Task I owned the ingestion utility and needed the load to finish reliably, without duplicates. [confirm: the deadline or downstream dependency that made it urgent]

Action I read the failures instead of re-running the job. They were HTTP 429 responses: the database was rate-limiting us because we sent more than the provisioned throughput. Retrying blindly would hammer it harder and could double-write. So I did three things. Retry with exponential backoff and a cap, so a 429 means wait, not crash. Idempotent upserts on a stable key, so a retry can never create a duplicate. And bulk execution instead of single writes, so far fewer round trips. I checked completion by reconciling source and target counts, not by the script exiting.

Result All 570K+ records landed, with a 10 to 20x throughput gain over the naive version. I used the same pattern on the historical load of 19 pricing and distribution feeds, which reached 17 of 19 through UAT during my time there.

Follow-up they will ask: "How would you parallelize without making the 429s worse?" Bounded concurrency sized to the provisioned throughput, with a shared backoff so every worker slows together.

S7. Saying "not yet" to a demo the VP liked (backbone, AI risk)

Fits: Have Backbone, Customer Obsession, Earn Trust, Success and Scale

Situation At a $1.3T asset manager, the sales team wrote meeting notes by hand, and turning them into CRM entries was slow and inconsistent. I built a Python prototype on Amazon Bedrock with Claude that turns free-form notes into structured, CRM-ready minutes, and demoed it live to the VP of Sales.

Task Give an honest feasibility call, not just a good demo.

Action I forced the output into the CRM's schema so a missing field would be obvious, and I tested on real messy notes rather than clean examples. That surfaced failure modes: ambiguous fields, and invented values when a note was thin. The demo went well. [confirm: what the VP said, and whether anyone wanted to go straight to production.] I said clearly it should not ship as it was, and offered a gated path: high-confidence extractions flow through, low-confidence ones go to a person to confirm, the write to the CRM respects user permissions, and an eval set built from real notes measures quality over time.

Result The concept was proven and the path to production was written down with its conditions, so nobody auto-wrote hallucinated values into a system of record at a regulated firm. [confirm: what happened next with the prototype.]

Follow-up they will ask: "How did you say no to someone who loved the demo?" I agreed with the goal, showed the failure cases on their own data, and gave a concrete route to yes with human review on low-confidence cases.

S8. Break it on purpose (GameDay to org standard)

Fits: Insist on the Highest Standards, Learn and Be Curious, Think Big, Hire and Develop the Best

Situation Prudential ran its first AWS Chaos Engineering GameDay, where teams inject failures into production-like systems and are scored on how they detect and recover. I had used chaos engineering before, with Gremlin at Renaissance Learning.

Task Compete with a 5-person cross-functional team, then make sure the lessons outlived the event.

Action During the event we stayed systematic as failures cascaded: stabilize first, then diagnose. The most useful person on the team was not the most senior; it was the engineer who kept reframing the problem when we hit dead ends, and I started building their path to senior from that day [confirm]. Afterwards I wrote the failure-injection approach and resilience patterns into standards and got them adopted beyond my own team. The rule I used: failure is survivable if you design the blast radius before you light the fuse.

Result We won the inaugural GameDay. The standards became the engineering org's baseline for resilience testing across multiple teams. I made sure the team's names were on the win, not only mine.

Follow-up they will ask: "How did you get other teams to adopt the standard?" [confirm: which teams, and the one objection you had to overcome]

S9. Growing engineers who ship without me (hiring and developing)

Fits: Hire and Develop the Best, Strive to be Earth's Best Employer, Ownership

Situation At Prudential I managed 15+ engineers across squads for three years, in a market where mid-level engineers were being poached.

Task I owned the full cycle: writing job descriptions, interviewing, headcount planning with HR, reviews, promotions, and exits.

Action I started promotion conversations about six months before the cycle: here is the gap to the next level, here is the evidence we need, here is a realistic timeline. Then I built the opportunities on purpose. If someone needed to show technical leadership, I gave them an architecture decision to own. If they needed visibility, I brought them into stakeholder meetings I usually attended alone. I kept a running note per person with three lines: what they are working on, what worries them, what they want to be doing in twelve months. I opened every one-on-one with "what do you need from me this week?"

Result I promoted 3+ engineers to senior and lead roles and kept retention above 90 percent [confirm 90 or 95] over three years. By the time I wrote each promotion document, nothing in it was a surprise to them.

Follow-up they will ask: "Tell me about one of those people." [confirm: pick one person, first name, their starting gap, the stretch assignment, and their role now]

S10. Killing the phone call (invent and simplify)

Fits: Invent and Simplify, Frugality, Deliver Results

Situation At Renaissance Learning, a K-12 EdTech company, engineering knowledge lived in a legacy knowledge base and in people's heads. Teams leaned on Zoom and phone calls to find answers, and support questions piled up.

Task I took on rebuilding how the team found information, so answers did not depend on catching the right person.

Action Instead of adding a new tool, I moved the knowledge base to Confluence, which the company already had, and set documentation standards: where things live, how pages are structured, how to write a note someone can use without a call. I coached teams on async-first, structured note-taking.

Result Retrieval time dropped 70 percent and support tickets dropped 90 percent. Calls and Zooms for routine questions fell by roughly 70 percent.

Follow-up they will ask: "How did you measure 70 percent?" [confirm: how retrieval time and ticket counts were measured, and over what period]

S11. Live data in demos without exposing a single customer record (security and risk)

Fits: Customer Obsession, Invent and Simplify, Success and Scale, Earn Trust

Situation At Vantage Commodities, a regulated trading firm, the sales team demoed our product with stale data. The real production data was covered by customer NDAs and could not be shown.

Task Give sales live, believable data for demos without any NDA-restricted raw record leaving its boundary.

Action I built an automated pipeline that produced production data extracts every 15 minutes into Tableau dashboards, shaped so the demo showed real patterns but never the restricted raw records. [confirm: exactly how, for example aggregation, masking, or a separate demo database with its own access.] It ran unattended, so nobody was tempted to copy data by hand.

Result Customers saw live data in demos, and the pipeline contributed to multiple sales conversions, with no NDA data exposed.

Follow-up they will ask: "Who signed off that it was safe?" [confirm]

S12. Bringing trading-floor controls into federal AI guidance

Fits: Success and Scale, Think Big, Earn Trust, Learn and Be Curious

Situation In 2026 I joined the community of interest drafting NIST's AI Risk Management Framework profile for critical infrastructure. Few members came from financial services, where firms already run hard controls on automated systems.

Task Bring that experience in as practical guidance, on volunteer time.

Action I proposed three things from finance. Hard limits outside the model, like the pre-trade checks under SEC Rule 15c3-5. Extending existing model-risk governance instead of creating a new AI office. And, working with a peer who proposed an aviation-style close-call system, non-punitive reporting of AI near-misses. I credited the peer's idea by name and drafted the part on sharing de-identified near-miss records between firms, citing ORX, where 80+ banks and insurers pool anonymized loss data.

Result The August draft added a pre-trade control example that was not in the July draft. The project lead credited my governance point by name in the community channel. Discussion Draft 3 (Sep 13, 2026) added a new near-miss reporting task, and its implementation on shared records closely matches the language I drafted.

Follow-up they will ask: "What is your part versus others'?" The separation of benefit-claiming from risk-reporting is the peer's idea. The de-identified shared-record piece is mine. It is a discussion draft and can still change.

S13. Making acquired engineers feel their work counted

Fits: Strive to be Earth's Best Employer, Earn Trust, Hire and Develop the Best

Situation Renaissance Learning acquired three companies while I was there: Flocabulary, Freckle, and FastBridge. Acquired engineers often leave in the first year.

Task I took on bridging the cultural and technical gaps for the acquired teams, beyond my engineering role.

Action I talked with the acquired engineers about what was pushing them out. The answer was not pay or tech stack. It was losing agency, and feeling their prior work was not valued. [confirm: one or two specific things you changed, for example pairing, bringing their patterns into our codebase, or giving them ownership of a component.] I carried the same finding to Prudential's M&A Inclusion Committee.

Result [confirm: any retention or engagement signal from the acquired teams.] The attrition finding became a standing input to how Prudential's committee approached integrations.

Follow-up they will ask: "How did you know agency was the cause?" They told me in one-on-ones, and the pattern held across three different acquisitions.

S14. Rostering for school districts (EdTech customer story)

Fits: Customer Obsession, Dive Deep, Deliver Results

Situation At Renaissance Learning I built student information system rostering: how a district's classes, teachers, and students flow into our learning products. Districts use different systems, and a bad sync means a teacher opens the app on the first day of school and their class is missing. [confirm scale: your inventory says 10,000+ school districts.]

Task Build rostering to the IMS Global OneRoster v1.1 standard in C# and .NET on AWS (EC2, RDS), and keep the APIs fast for large districts.

Action I followed the OneRoster spec so districts could plug in without custom work, and designed the sync pipelines for district integrations. When large payloads made read endpoints slow, I added pagination and database indexing on the high-volume queries.

Result Rostering ran at national K-12 scale and API response times improved under large payloads. [confirm: a before and after number for the slow endpoints.]

Follow-up they will ask: "Why follow a standard instead of custom integrations?" One standard means each new district is configuration, not a project. That is the same reason this team pushes ISVs toward repeatable patterns.

Story map

Rows are principles, columns are stories. A check means the story is a good fit. Count the checks per row to see where you are strong and where you are thin.

PrincipleS1S2S3S4S5S6S7S8S9S10S11S12S13S14
Customer Obsession✓✓✓✓✓
Earn Trust✓✓✓✓✓✓✓
Dive Deep✓✓✓✓
Ownership✓✓✓
Deliver Results✓✓✓✓✓✓
Invent and Simplify✓✓✓✓
Have Backbone✓
Learn and Be Curious✓✓✓
Bias for Action✓
Are Right, A Lot✓✓✓
Hire and Develop✓✓✓
Highest Standards✓✓✓
Think Big✓✓✓
Frugality✓✓✓
Earth's Best Employer✓✓
Success and Scale✓✓✓

Story list: S1 conference app miss, S2 Vantage cutover, S3 Kendra, S4 Advisor Portal, S5 eval, S6 Cosmos 429s, S7 Bedrock CRM, S8 GameDay, S9 growing engineers, S10 Confluence, S11 Tableau demo data, S12 NIST, S13 M&A inclusion, S14 rostering.

Gaps, said plainly

  1. Have Backbone is thin. One story, and the vault does not record the VP's reaction. Candidates to build [confirm each is real before using]: the Vantage decision to keep C++ on hot paths, if someone senior wanted a full migration; or the high performer at Prudential who was dismissive of junior engineers, from your June 2026 prep notes, where you named the pattern and then reduced their scope.
  2. Bias for Action is thin. The Cosmos story is more Dive Deep than speed. Candidates [confirm]: ParisTransitHelper, which you built right after the Bus 63 ticket failure in Paris; or the APIGEE to Kong migration at PGIM, where you ran cross-team whiteboarding to produce the API inventory and target design.
  3. No external-customer SA story yet. This role is about convincing an ISV's engineering team. Your stories are mostly internal. Candidates [confirm]: your ITG pre-sales work (technical scoping, effort estimates, delivery approach for prospective clients); the national marketing review of the conference app, where you had to satisfy another organization's brand and accessibility bar; or the second chapter's request, if that conversation has moved forward.
  4. Earth's Best Employer has thin actions. S13 needs one or two concrete things you did. The Prudential reorg answer in your prep notes could help if it happened as written [confirm].
  5. Ownership and Are Right, A Lot are fine but reuse Vantage. Keep Vantage for one interviewer only, and use S1 and S3 for the others.

SA-flavored behavioral questions

These are phrased the way Amazon interviewers phrase them for a Solutions Architect. Each answer is an outline pointing to your stories. Say it in your own words.

Tell me about a time you had to convince a customer's engineering team to change their technical approach.

Closest real story: S3, Kendra. The "customer" was the operations team and the risk and legal reviewers. Open with the result (20+ steps cut, production for 35,000+ advisors). The convincing: I came with the compliance documentation and rollback criteria before they asked, and I tuned relevance against their real queries so they saw their own work get easier. Say plainly it was an internal customer, then add how you would do the same with an ISV. [confirm an external version from ITG pre-sales]

Tell me about a time a customer wanted to go to production faster than was safe.

S7, the Bedrock CRM prototype. Show the failure cases you found on real notes, the schema constraint, and the gated path: human review for low-confidence extractions, permission-aware CRM writes, and an eval set. Close with the principle: I give a no together with a route to yes.

Tell me about the most complex technical problem you solved for a customer. Go as deep as you can.

S6, Cosmos 429s, then let them peel. Be ready to go down to: what a 429 means, provisioned throughput, exponential backoff with a cap, idempotent upsert on a stable key, bulk execution, count reconciliation. If they want architecture scale instead, switch to S2, Vantage parallel run.

Tell me about a time you made a technical recommendation that turned out to be wrong.

S1 is your cleanest mistake, but it is an omission, not a wrong recommendation. Tell it, then add S5's smaller one: your first cross-language metric measured agreement, not correctness, and you caught it yourself and changed it. Two honest misses, both with the fix.

Tell me about a time you had to learn a new technology quickly to be useful to a customer.

Your AI rebuild over the last year: on-device RAG with an eval (S5), a Claude assistant with OAuth and streaming, a multi-agent recipe system, Claude Code tooling with MCP. Method: read the existing pattern, ship something small that works, then harden it. Result: several apps in production built solo. Tie it to Bedrock: the same method is how you would ramp on a new Bedrock feature for a customer.

Tell me about a time you built something that was reused far beyond its original purpose.

S1: the conference app as one typed data file, reuse designed before anyone asked, second chapter asking within five days. Also the retrieval engine reused across three constitution projects (US Declaration, India with 609 passages, Liechtenstein with 122 articles). This is Think Big and Invent and Simplify together, and it matches the SA job of turning one ISV's win into a pattern.

Tell me about a time you had to explain a complex technical trade-off to a non-technical executive.

S2: keeping C++ on hot paths, explained to traders and the CTO as protecting execution speed where milliseconds move money. Or S7 to the VP of Sales. Make the move explicit: I translate to revenue, risk, and time, never the stack.

Tell me about a time you raised a security or compliance risk that others had missed.

S11 (NDA-restricted data kept out of sales demos by design) or S7 (no unattended AI writes into a system of record). For public sector, add S12: hard limits outside the model and near-miss reporting. Name the harm, the control, and that you designed it before anyone asked.

Tell me about a time you disagreed with a senior leader and how it was resolved.

S7 today. Build the second backbone story from the gaps list before the loop. Structure: what they wanted, the data you brought, how you raised it (privately, once, in writing), what was decided, and how you committed. If you lost, say so; losing and committing scores.

Tell me about a time you delivered a project under a hard constraint you could not change.

S1: fixed conference date, no budget, unreliable venue wifi. Or S2: the desk could not stop trading. Name the constraint first, then the design it forced (static export and offline caching; parallel run with rollback).

Tell me about a time you used data to change someone's mind.

S5: the numbers themselves (0.55 to 0.77 for real answers, 0.35 cap for off-topic) are what make the case that the tool abstains. Or S3's before and after in the workflow. Say what they believed before, the data, and what they decided after.

Tell me about a time you helped a customer reduce cost.

S5 (retrieval-only, zero per-query inference cost, no backend), S1 (static CDN, zero hosting cost), or the Vantage scrapers that replaced paid market-data subscriptions. Then the SA link: for Beacon, right-size the model per task and cache what repeats.

Tell me about a time you mentored someone outside your reporting line.

NYU Tandon capstone mentoring (Fall 2025 and Spring 2026 cohorts, including a Capgemini-sponsored ML project on insurance pricing), or the Claude Code coaching you run for engineers and analysts. [confirm one named mentee and one concrete outcome]

Tell me about a time you made a decision with incomplete data.

Build this from the Bias for Action candidates. Until then, S3: you chose a managed service over a custom pipeline before you could compare both in production, because it was a two-way door and the deadline was real. Say "two-way door" out loud.

Why do you want to be an SA on this team, and what will you do in the first 90 days?

Not a Leadership Principle question, but it comes up. Your reasons: you have built for both sides of this team's customers (K12 at Renaissance, regulated systems at Prudential and Vantage); you have shipped AWS AI to production (Kendra, Bedrock); and you like enabling other engineers. First 90 days: listen to two or three ISVs, pick one with a clear GenAI use case, ship a working pattern with guardrails, and write it up so the next ISV starts from it.

Traps