Why This Architecture Matters

Most enterprise RAG demos die in the prototype phase. The gap between "we embedded our PDFs in a vector DB" and "we ship to thousands of enterprise tenants with sub-second TTFB" is where teams burn quarters.

Gallup's case is worth studying because they didn't invent anything exotic. They picked boring, managed primitives — Amazon Bedrock, Knowledge Bases, Kendra, Lambda, ElastiCache Serverless — and wired them together with discipline. The result: production in weeks, no dedicated MLOps team, and prompt volume up ~7x since June 2024.

This post breaks down the reference architecture, the non-obvious design decisions, and where this pattern will bite you if you copy it blindly.

근거자료: AWS Architecture Blog — Gallup delivers real-time workplace coaching

Enterprise AI assistant chat interface powered by Amazon Bedrock delivering personalized workplace coaching Software Concept Art

The Reference Architecture, Decoded

At its core, this is a dual-retrieval RAG pipeline behind a streaming API, with guardrails and observability bolted on. Let's walk the request path.

1. Ingestion: Two Sources, Two Indexes

# Simplified ingestion flow (pseudo-code for clarity)
# 수집 파이프라인: 정적 리서치 + 실시간 웹 크롤링

sources = {
    "historical": "s3://gallup-research-archive/",   # 90 years of studies
    "live": "https://www.gallup.com/insights/",       # daily crawl
}

# Static archive → Bedrock Knowledge Bases (managed vector store)
bedrock_kb.ingest(sources["historical"])

# Live site → Kendra (keyword + semantic hybrid)
kendra.index_crawl(sources["live"], refresh="daily")

Why two indexes instead of one? Because historical research and current publications have different retrieval characteristics. The archive is stable, dense, and benefits from embedding-based similarity. The live site changes daily and benefits from Kendra's hybrid keyword+semantic ranking. This is a pattern worth stealing.

2. Query Time: Score, Filter, Consolidate

# 요청 처리: 두 인덱스에서 검색 → 신뢰도 기반 필터링 → 통합
def answer(user_query: str, user_context: dict):
    # Retrieve from both sources in parallel
    kb_hits = bedrock_kb.retrieve(user_query, top_k=10)
    kendra_hits = kendra.query(user_query, top_k=10)

    # Score by confidence threshold — this is the key filter
    candidates = [
        h for h in (kb_hits + kendra_hits)
        if h.confidence >= CONFIDENCE_THRESHOLD
    ]

    # Consolidate & deduplicate before hitting the LLM
    context = consolidate(candidates)

    # Stream via Claude on Bedrock with guardrails applied
    return bedrock.invoke_stream(
        model="anthropic.claude",
        prompt=build_prompt(user_query, context, user_context),
        guardrail_id=GUARDRAIL_ID,
    )

The CONFIDENCE_THRESHOLD filter is doing a lot of work here. Without it, you flood the context window with marginally relevant chunks and the model hallucinates connections. With it, you trade recall for precision — which is the right tradeoff for a coaching assistant where a wrong answer damages trust.

3. The Serverless Backbone

LayerServiceWhy
ComputeAWS Lambda + FastAPIReal-time streaming without managing servers
Session cacheElastiCache ServerlessSub-ms conversation history retrieval
Durable storeRDS for MySQLSystem of record for prompts/responses/citations
Fast KVDynamoDBPer-user contextual insights
MetricsData Firehose → S3Token I/O, TTFB, stop reasons for cost tuning
ConfigSSM Parameter StoreChange model/policy settings without redeploy

The SSM Parameter Store as config hub is the underrated move. Being able to swap model versions or tighten guardrails without a deploy is what lets a small team operate this at scale.

Serverless AWS architecture diagram with Lambda, Bedrock Knowledge Bases, and Kendra for RAG pipeline Programming Illustration

What Actually Moved the Metrics

Since June 2024, per Gallup's published numbers:

MetricGrowth
Prompts~7x
Conversations~4.5x
Active users~5.5x
Prompts per conversation~55% ↑
TTFB (streaming)Sub-second
Session retrievalSub-millisecond

The 55% increase in prompts-per-conversation is the most interesting number. It signals users aren't just kicking the tires — they're having multi-turn coaching sessions. That's a product-market signal, not just a tech signal.

Where This Pattern Will Bite You

A few honest caveats before you copy this architecture:

  1. Dual-index maintenance is real work. Keeping Kendra's crawl and Knowledge Bases' embeddings in sync is not free. If your content doesn't have a clear static/dynamic split, use one index.
  2. Guardrails mid-stream are tricky. Intervening after tokens have already been streamed to the client means the UI must handle retraction gracefully. Design your frontend for this.
  3. ElastiCache Serverless pricing scales with usage. Sub-ms is great until your conversation volume 10x's. Model your costs before committing.
  4. Confidence thresholds are model-dependent. A threshold tuned for Claude 3.5 will misbehave when you swap to a newer model. Version your thresholds alongside your model IDs in SSM.
  5. Kendra is not cheap. For smaller corpora, pgvector or OpenSearch Serverless may be more economical.

The Roadmap Signal: Agents Next

Gallup is building on Amazon Bedrock AgentCore — moving from a user-facing assistant to a programmatic layer that other systems can call. This is the pattern to watch in 2025: RAG assistants are becoming the front door, and agentic orchestration is becoming the back office.

If you want to see how a different hyperscaler handles the routing layer for ML inference at extreme scale, the Netflix model-serving architecture is a great counterpoint to study: How Netflix Routes 1M+ ML Inference Requests Per Second.

And if you're working on the frontend side of these AI products, the weekly CSS roundup covering modern techniques like SVG favicons and anchor-interpolated morphing is worth a look: CSS Weekly Roundup.

Cloud infrastructure dashboard showing token metrics, TTFB latency, and Firehose streaming to S3 Coding Session Visual

What to Take Away

Gallup's build is not a research project — it's an operating pattern. The three decisions that matter most:

  • Managed over custom. Bedrock + Knowledge Bases + Kendra meant no MLOps hire. If your team is small, buy the primitives.
  • Streaming UX is non-negotiable. Sub-second TTFB is the difference between a tool people use and a tool people abandon.
  • Ground in verified data, not vibes. The dual-index approach keeps responses anchored to research instead of letting the LLM freestyle.

Next Steps

If you're building something similar:

  1. Start with Amazon Bedrock Knowledge Bases and a single corpus. Get the RAG loop working end-to-end before adding Kendra.
  2. Instrument TTFB and tokens-per-response from day one. Cost surprises kill these projects.
  3. Read AWS's prescriptive guidance on writing best practices for RAG applications to avoid the common retrieval pitfalls.
  4. When you're ready to go beyond a chat interface, look at Bedrock AgentCore for exposing your knowledge programmatically.

The bar for enterprise AI has shifted. It's no longer "can you build a RAG demo?" — it's "can you ship a grounded, observable, cost-controlled assistant that thousands of users trust?" Gallup's stack is a solid blueprint for answering yes.

This content was drafted using AI tools based on reliable sources, and has been reviewed by our editorial team before publication. It is not intended to replace professional advice.