RAG on a Raspberry Pi: Cited Answers From a $100 Board

Stock research is a hallucination minefield: a confident answer with a fabricated citation is worse than no answer. VietProStocks answers questions about Vietnamese equities with cited sources while running entirely on a Raspberry Pi 4 8GB. This is how the pipeline works and what it costs.
The boring pipeline wins
Every stage is deliberately simple, because simple stages are debuggable without a GPU:
Chunk → embed → index → rerank → answer with citations.
# ingestion: 500-token chunks with 100-token overlap
for chunk in chunk_document(doc, size=500, overlap=100):
vector = embed(chunk.text) # all-MiniLM-L6-v2
index.add(vector, metadata=chunk.meta)
700 token chunks failed (answers lost surrounding context); 200-token chunks flooded retrieval with fragments. 500/100 was measured, not guessed.
Rerank with a trust signal, not just similarity
Retrieval pulls 6–12 candidates; the final prompt gets 3–5. The rerank formula is explicit:
score = semantic_similarity
+ recency_decay(published_at)
+ source_trust[source]
source_trust is configured per source. A scraped forum and an official filing do not enter the prompt as equals, and recency alone would have surfaced low-trust noise from a thin Vietnamese news ecosystem.
Validate every citation before display
The model is told to cite chunk sources — and then the server ignores the model's word for it. Each citation URL is checked against the indexed corpus; anything that does not resolve is stripped before the answer renders. Hallucinated sources never reach the UI.
answer = llm.generate(prompt, context=chunks)
answer.citations = [c for c in answer.citations if c.url in index.known_urls]
This single gate did more for trustworthiness than any prompt engineering.
The hardware truth
A Q4-quantized 7B model on a Pi answers in seconds to minutes. That is a product decision, not a bug:
- Interactive chat — use a nearby machine (or the OpenAI-compatible backend) when latency matters.
- Scheduled research and alerts — the Pi is perfect: it runs overnight, costs nothing, and never leaks your query history.
Embedding on CPU lands at roughly 1–3 seconds per document, which is fine for background ingestion. The FAISS index grows to hundreds of megabytes; we keep IndexFlatIP for accuracy and documented IndexHNSWFlat as the upgrade path when the corpus doubles.
Operations are part of the ML system
An index is a build artifact. Losing it should cost time, not data:
- systemd units, Cloudflare Tunnel for HTTPS, no Kubernetes for one board
- daily database backups, weekly index backups, keep seven generations
- a documented rebuild path from the source corpus
Three lessons
- Retrieval quality beats model size. Most bad answers were bad context, not a weak model.
- Latency is a product decision. The same engine serves chat and batch jobs if the product admits which is which.
- Every citation is a claim. Validate claims mechanically, or publish hallucinations.
Zero per-query cost, zero data leaving the device, and answers you can trace to a URL. For a $100 board, that is a good trade.
