During Juno RAG’s build, it became clear that embeddings alone don’t guarantee trustworthy answer...During Juno RAG’s build, it became clear that embeddings alone don’t guarantee trustworthy answer...
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
During Juno RAG’s build, it became clear that embeddings alone don’t guarantee trustworthy answers. High cosine similarity can still surface partially relevant or context-drifted chunks.
To address this, the pipeline was engineered with:
• Hybrid search (semantic + BM25) to balance meaning and exact match
• Cross-encoder re-ranking to evaluate true query-document relevance
• Context-aware chunking to preserve semantic continuity
• Mandatory source citation to enforce answer traceability
This shifted the system from probabilistic guessing to evidence-backed response generation.
Because in production RAG systems,
retrieval precision > generation fluency
Then: 1 month, 1 template.
Now: 1 week, 15 templates.
Same designer, different workflow.
I wrote up how I built no-code.supply: fifteen website templates, each with its own brand, in HTML and React, and some in Framer too. Claude Code did most of the typing. I did the directing.
An AI-built app can work perfectly while any stranger can read its customers' data. Nothing on screen shows it.
Ask whoever built it: which tables hold user data, and what stops a logged-out stranger from reading each one? A good answer is a list. "It should be fine" means nobody has checked.
Andrés, this is the check most founders never run, because nothing breaks until someone looks. Asking for the list of tables and who can read each one is a test any owner can do without writing code. Row level access belongs on the launch checklist, not the cleanup list.
Completed "Building and Evaluating Advanced RAG" by DeepLearning.AI (TruEra, LlamaIndex).
The most useful part wasn't building another RAG pipeline, it was learning how to actually evaluate one. The RAG Triad (answer relevance, context relevance, groundedness) gives you a way to catch exactly where a system is failing: bad retrieval vs. the model making things up vs. just answering the wrong question.
Also went through sentence-window and auto-merging retrieval, two ways to fix the classic RAG trade-off between precise matching and having enough context to actually answer well.