Why This Architecture Matters
Most enterprise RAG demos die in the prototype phase. The gap between "we embedded our PDFs in a vector DB" and "we ship to thousands of enterprise tenants with sub-second TTFB" is where teams burn quarters.
Gallup's case is worth studying because they didn't invent anything exotic. They picked boring, managed primitives — Amazon Bedrock, Knowledge Bases, Kendra, Lambda, ElastiCache Serverless — and wired them together with discipline. The result: production in weeks, no dedicated MLOps team, and prompt volume up ~7x since June 2024.
This post breaks down the reference architecture, the non-obvious design decisions, and where this pattern will bite you if you copy it blindly.
근거자료: AWS Architecture Blog — Gallup delivers real-time workplace coaching

The Reference Architecture, Decoded
At its core, this is a dual-retrieval RAG pipeline behind a streaming API, with guardrails and observability bolted on. Let's walk the request path.
1. Ingestion: Two Sources, Two Indexes
# Simplified ingestion flow (pseudo-code for clarity)
# 수집 파이프라인: 정적 리서치 + 실시간 웹 크롤링
sources = {
"historical": "s3://gallup-research-archive/", # 90 years of studies
"live": "https://www.gallup.com/insights/", # daily crawl
}
# Static archive → Bedrock Knowledge Bases (managed vector store)
bedrock_kb.ingest(sources["historical"])
# Live site → Kendra (keyword + semantic hybrid)
kendra.index_crawl(sources["live"], refresh="daily")
Why two indexes instead of one? Because historical research and current publications have different retrieval characteristics. The archive is stable, dense, and benefits from embedding-based similarity. The live site changes daily and benefits from Kendra's hybrid keyword+semantic ranking. This is a pattern worth stealing.
2. Query Time: Score, Filter, Consolidate
# 요청 처리: 두 인덱스에서 검색 → 신뢰도 기반 필터링 → 통합
def answer(user_query: str, user_context: dict):
# Retrieve from both sources in parallel
kb_hits = bedrock_kb.retrieve(user_query, top_k=10)
kendra_hits = kendra.query(user_query, top_k=10)
# Score by confidence threshold — this is the key filter
candidates = [
h for h in (kb_hits + kendra_hits)
if h.confidence >= CONFIDENCE_THRESHOLD
]
# Consolidate & deduplicate before hitting the LLM
context = consolidate(candidates)
# Stream via Claude on Bedrock with guardrails applied
return bedrock.invoke_stream(
model="anthropic.claude",
prompt=build_prompt(user_query, context, user_context),
guardrail_id=GUARDRAIL_ID,
)
The CONFIDENCE_THRESHOLD filter is doing a lot of work here. Without it, you flood the context window with marginally relevant chunks and the model hallucinates connections. With it, you trade recall for precision — which is the right tradeoff for a coaching assistant where a wrong answer damages trust.
3. The Serverless Backbone
| Layer | Service | Why |
|---|---|---|
| Compute | AWS Lambda + FastAPI | Real-time streaming without managing servers |
| Session cache | ElastiCache Serverless | Sub-ms conversation history retrieval |
| Durable store | RDS for MySQL | System of record for prompts/responses/citations |
| Fast KV | DynamoDB | Per-user contextual insights |
| Metrics | Data Firehose → S3 | Token I/O, TTFB, stop reasons for cost tuning |
| Config | SSM Parameter Store | Change model/policy settings without redeploy |
The SSM Parameter Store as config hub is the underrated move. Being able to swap model versions or tighten guardrails without a deploy is what lets a small team operate this at scale.

What Actually Moved the Metrics
Since June 2024, per Gallup's published numbers:
| Metric | Growth |
|---|---|
| Prompts | ~7x |
| Conversations | ~4.5x |
| Active users | ~5.5x |
| Prompts per conversation | ~55% ↑ |
| TTFB (streaming) | Sub-second |
| Session retrieval | Sub-millisecond |
The 55% increase in prompts-per-conversation is the most interesting number. It signals users aren't just kicking the tires — they're having multi-turn coaching sessions. That's a product-market signal, not just a tech signal.
Where This Pattern Will Bite You
A few honest caveats before you copy this architecture:
- Dual-index maintenance is real work. Keeping Kendra's crawl and Knowledge Bases' embeddings in sync is not free. If your content doesn't have a clear static/dynamic split, use one index.
- Guardrails mid-stream are tricky. Intervening after tokens have already been streamed to the client means the UI must handle retraction gracefully. Design your frontend for this.
- ElastiCache Serverless pricing scales with usage. Sub-ms is great until your conversation volume 10x's. Model your costs before committing.
- Confidence thresholds are model-dependent. A threshold tuned for Claude 3.5 will misbehave when you swap to a newer model. Version your thresholds alongside your model IDs in SSM.
- Kendra is not cheap. For smaller corpora, pgvector or OpenSearch Serverless may be more economical.
The Roadmap Signal: Agents Next
Gallup is building on Amazon Bedrock AgentCore — moving from a user-facing assistant to a programmatic layer that other systems can call. This is the pattern to watch in 2025: RAG assistants are becoming the front door, and agentic orchestration is becoming the back office.
If you want to see how a different hyperscaler handles the routing layer for ML inference at extreme scale, the Netflix model-serving architecture is a great counterpoint to study: How Netflix Routes 1M+ ML Inference Requests Per Second.
And if you're working on the frontend side of these AI products, the weekly CSS roundup covering modern techniques like SVG favicons and anchor-interpolated morphing is worth a look: CSS Weekly Roundup.
![]()
What to Take Away
Gallup's build is not a research project — it's an operating pattern. The three decisions that matter most:
- Managed over custom. Bedrock + Knowledge Bases + Kendra meant no MLOps hire. If your team is small, buy the primitives.
- Streaming UX is non-negotiable. Sub-second TTFB is the difference between a tool people use and a tool people abandon.
- Ground in verified data, not vibes. The dual-index approach keeps responses anchored to research instead of letting the LLM freestyle.
Next Steps
If you're building something similar:
- Start with Amazon Bedrock Knowledge Bases and a single corpus. Get the RAG loop working end-to-end before adding Kendra.
- Instrument TTFB and tokens-per-response from day one. Cost surprises kill these projects.
- Read AWS's prescriptive guidance on writing best practices for RAG applications to avoid the common retrieval pitfalls.
- When you're ready to go beyond a chat interface, look at Bedrock AgentCore for exposing your knowledge programmatically.
The bar for enterprise AI has shifted. It's no longer "can you build a RAG demo?" — it's "can you ship a grounded, observable, cost-controlled assistant that thousands of users trust?" Gallup's stack is a solid blueprint for answering yes.