Archived copy of an article by Omer Sen, originally published on LinkedIn on 2026-06-29.
Original: https://www.linkedin.com/pulse/wired-langfuse-observability-rag-pipeline-omer-sen-xthve/ · ← back to faruk.net
Original: https://www.linkedin.com/pulse/wired-langfuse-observability-rag-pipeline-omer-sen-xthve/ · ← back to faruk.net
Wired Langfuse observability into a RAG pipeline
Wired Langfuse observability into a RAG pipeline — here's what you actually see
I already built a production RAG pipeline on Azure (AI Search + OpenAI + Terraform). What I hadn't done was add proper LLM tracing to it.
This project wraps every stage of the pipeline with Langfuse v4's @observe decorator:
- Retrieval — docs returned, top similarity score, latency
- Generation — model, prompt, completion, token counts
- Quality scores — context relevance + answer faithfulness per trace
@observe(name="rag-query")
def query(self, question: str) -> dict:
retrieved_docs = self._retrieve(question)
answer = self._generate(question, retrieved_docs)
self.langfuse.score_current_trace(name="context_relevance", value=score)
The gap between "RAG demo" and "RAG in production" is almost entirely observability. Without traces you can't tell if retrieval is failing, if the model is hallucinating, or where the latency is.
Docker:
docker run --rm -e ANTHROPIC_API_KEY=key \
-e LANGFUSE_PUBLIC_KEY=pk-lf-... \
omerfsen/llm-tracing-langfuse:latest
#Langfuse #LLMOps #RAG #AIEngineering #Python #OpenSource