METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation (ACM SOSP25) Siddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du, Shaoting Feng, Ganesh Ananthanarayanan, Ravi Netravali, Junchen Jiang RAG systems are ubiquitous in todays LLM
"METIS: Fast Quality-Aware RAG Systems with" is the property of its rightful owner. Permission is granted to
download and print the materials on this website for personal, non-commercial use only, and to display it
on your personal computer provided you do not modify the materials and that you retain all copyright
notices contained in the materials. By downloading content from our website, you accept the terms of this
agreement.
Presentation Transcript
01
METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation(ACM SOSP’25) Siddhant Ray, Rui Pan, Zhuohan Gu, Kuntai Du, Shaoting Feng, Ganesh Ananthanarayanan, Ravi Netravali, Junchen Jiang<br>
02
RAG systems are ubiquitous in today’sLLM applications RAG (Retrieval Augmented Generation)
retrieves relevant context for the query based on vector similarity search
combines the chunks with the query to generate the response
Two directions of current work:
Direction 1 : Focus only on generation quality (AdaptiveRAG ACL’24)
Direction 2 : Focus only on low latency/high throughput (Parrot OSDI’24)
Focus of this talk:
Can we jointly optimize the quality-efficiency tradeoff for RAG serving? 2<br>
03
Configuration knobs RAG queries need RAG systems inherently expose multiple configuration knobs 3<br>
04
Overview on synthesis_method Chunk 1 Chunk 2 Chunk 3 LLM Final Answer Chunk 1 Chunk 2 Chunk 3 Final Answer 1
Confidence : 80% Final Answer 2
Confidence : 99% Final Answer 3
Confidence : 90% Chunk 1 Chunk 2 Chunk 3 S1 S2 S3 Final Answer (a) Stuff (b) Map Rerank (c) Map Reduce LLM LLM LLM 4<br>
05
Optimizing RAG’s tradeoff space is challenging RAG queries are underspecified
Query only has natural language, no operators
Heterogeneity in RAG queries leads to per-query quality-latency trade-offs
Add more chunks -> (often) higher quality and latency
Add more chunks -> (often) higher latency but not higher quality
RAG serving greatly benefits from joint quality-latency optimization 5<br>
06
Change: synthesis method from map_rerank
(circle) , stuff (plus) and map_reduce (square) Q1 Q2 Q3 Q1: In what county was William W. Blair born?
Q2: Are Alison Skipper, Diane Gilliam Fisher, and Rachel McAdams from the same country?
Q3: When and why did the Voyager 1, the spacecraft that detected storms on Neptune, leave our solar system? RAG quality-latency trade-offs vary at thequery level! 6<br>
07
Per-query configuration adaptation achieves better quality-latency trade-offs Pareto Boundary of fixed configuration
with vLLM Pareto Boundary of fixed configuration
with vLLM Per-Query
Configuration Per-Query
Configuration 7<br>
08
METIS : Per-query RAG configuration adaptation RAG Queries Retriever RAG Synthesis Generated
Output Chosen Config 8 First RAG controller on top of traditional RAG pipeline! Per-query LLM profiler reduces the space of RAG configurations
Schedules RAG queries jointly considering quality and latency<br>
09
Online profiling RAG queries reduces thesize of the tradeoff space Query Profiler
( LLM ) Input Prompt Query Rule-based
Mapping Resource
AwareScheduler Query
Profile Pruned
Config
Space E.g., Is query-complexity high/low? E.g., If query-complexity == high ->
Use stuff / map_reduce Range of useful configs for given
query 9<br>
10
System Resource-Aware Adaptation forJoint Scheduling (a) Baseline separates configuration selection and scheduling (b) METIS performs configuration selection and scheduling jointly 10<br>
11
Detecting if the profiling fails Above threshold -
98% good profiles Above threshold -
96% good profiles 7% below threshold -90% bad profiles 7% below threshold -85% bad profiles 90% Threshold 11<br>
More results in the paper Cost analysis of the profiler
Showing minimal latency overhead of the profiler
Sensitivity analysis for serving model and profiler
Feedback-based system improvement over time 16<br>
17
Conclusion First system to focus on optimizing the tradeoffs between response delay and generation quality for RAG workloads
Schedules RAG queries and adapts key RAG configuration knobs on a per-query basis.
Result :
1.64x-2.54x delay reduction at same response quality!
1.8x-4.5x higher throughput vs closest-quality fixed configurations! 17<br>