Technical Reference
v2.0.0How Lingua Rakyat Works
Architecture reference — ingestion pipeline, Q&A system, tech stack, API, and eval metrics.
Project Overview
- Government PDFs written in legalese — not plain language
- Long documents with no searchable Q&A interface
- No equal access across Malay, English, and Chinese speakers
- Eligibility criteria buried in paragraphs
- Upload any government PDF — queryable instantly
- Ask in Malay, English, or Mandarin — answer in same language
- Every answer grounded in document text with page citations
- Evidence guard refuses to hallucinate — shows confidence
RAG in one sentence: Instead of an LLM guessing from training memory, Lingua Rakyat first searches the actual document for relevant passages, then passes only those passages to the LLM — so the answer is always traceable to a real source.
Ingestion Pipeline
Phase 1Triggered once when a PDF is uploaded. Runs offline — users can start asking questions as soon as ingestion completes.
- Checks file type, encryption status, and page count (1–500 pages)
- Extracts text with pypdf — pages with < 50 chars flagged for OCR
- Rejects PDFs where > 80% of pages have no extractable text
- pypdf extracts text page-by-page for text-based PDFs
- PyMuPDF renders low-text pages to bitmap at 2× scale
- Tesseract OCR reads bitmaps — supports eng + msa + chi_sim
- Regex detects section headers: Bahagian, Seksyen, Section, numbered (1.2.3), ALLCAPS
- Each section becomes its own chunk — preserves document structure
- Long sections split at 360-word target, 45-word overlap, min 20 words/chunk
- All chunks embedded in batch via Cohere embed-multilingual-v3.0
- Produces 1024-dim vectors per chunk — language-agnostic representation
- input_type = search_document (optimised for retrieval, not clustering)
- Vectors + metadata → Pinecone under document namespace
- Raw PDF file → Supabase Storage bucket
- Document record (name, chunk count, timestamps) → Supabase PostgreSQL
Q&A Pipeline
Phase 2Triggered per question. Streams tokens to the UI as they're generated — first token appears within ~600ms on a warm cache.
- Keyword matching first: Malay words (nak, boleh, saya, mohon) with word-boundary regex
- CJK character ratio: if > 15% of chars are CJK, classify as zh-cn
- langdetect library as fallback — maps 20+ dialects to en / ms / zh-cn
- Expands 1 question into up to 4 variants: original + paraphrase + translations
- Translations generated into all other supported languages
- Improves retrieval recall — same chunks found regardless of query language
- All variants embedded in a single Cohere API call (input_type = search_query)
- Each variant queries Pinecone — top-k chunks per variant
- Results merged, deduplicated by chunk ID, sorted by composite score
- Cohere rerank-multilingual-v3.0 re-scores candidates reading query + document together
- Final score = vector_score × 0.35 + rerank_score × 0.65
- Higher precision than cosine similarity alone — cross-encoder architecture
- Top chunk score ≥ 0.50 → strong evidence → standard QA prompt
- Score 0.12–0.49 → cautious mode → answer with caveats
- Score < 0.12 → hard refusal → canned message, no hallucination
- Context-aware prompt built in detected language (en/ms/zh-cn)
- Last 3 conversation turns injected for follow-up awareness
- Groq LLaMA 3.3 70B streams answer in 3–5 bullet points token-by-token
- Faithfulness: answer passed back through Cohere reranker vs source chunks
- LLaMA 3.1 8B generates 3 follow-up question suggestions (non-blocking)
- Result cached (LRU, 200-entry max) — cache bypassed when chat history present
API Reference
All endpoints rate-limited per IP via SlowAPI. Interactive docs at /docs (Swagger UI).
| Method | Path | Description |
|---|---|---|
| POST | /api/documents/upload | Upload + validate + ingest PDF into Pinecone |
| GET | /api/documents/ | List all documents with metadata |
| DELETE | /api/documents/{id} | Delete PDF from Supabase + vectors from Pinecone |
| POST | /api/chat/ask-stream | Streaming Q&A via Server-Sent Events |
| POST | /api/chat/ask | Buffered Q&A (non-streaming) |
| GET | /api/chat/history | Load chat history for a document |
| DELETE | /api/chat/history/{doc_id} | Clear chat history for a document |
| POST | /api/eval/run-test-suite-stream | Run benchmark suite with live streaming results |
| GET | /api/eval/report | Get aggregated ROUGE/BLEU/latency metrics |
| GET | /api/eval/data-quality | Chunk quality stats per document |
| POST | /api/voice/transcribe | STT — WebM/Opus audio → text via Groq Whisper |
| POST | /api/voice/tts | TTS — text → MP3 audio via ElevenLabs |
| POST | /api/feedback | Submit thumbs up/down → persisted in Supabase |
Tech Stack
Backend
What: Python async web framework. Auto-generates Swagger docs at /docs. High throughput via async I/O.
Why: Serves all endpoints — chat, documents, eval, voice — with automatic OpenAPI documentation for judges.
What: Meta's open-source LLM (70B params) running on Groq's custom LPU inference hardware. Extremely fast token generation.
Why: Generates grounded, source-cited answers in Malay, English, or Chinese from retrieved document context.
What: Smaller, faster variant of LLaMA for low-latency tasks where full 70B is overkill.
Why: Generates follow-up question suggestions after each answer without blocking the main response stream.
What: Embedding model supporting 100+ languages. Converts text into 1024-dim vectors capturing semantic meaning.
Why: Embeds both document chunks (at ingestion) and queries (at retrieval) enabling cross-lingual semantic search.
What: Cross-encoder model that re-scores retrieved candidates by reading query + document together.
Why: Re-ranks vector search results in context — higher precision than pure cosine similarity. Also computes faithfulness.
What: Managed cloud vector database. Stores and searches high-dimensional embeddings at scale.
Why: Stores all document chunk vectors. Each document gets its own namespace for isolated retrieval.
What: Open-source Firebase alternative — PostgreSQL + file storage + real-time subscriptions.
Why: Stores raw PDFs (Storage bucket), document metadata, chat history per user/document, feedback thumbs.
What: pypdf: pure-Python text extraction. PyMuPDF (fitz): C-based PDF renderer for rasterising pages to images.
Why: pypdf handles text-based PDFs. PyMuPDF renders scanned pages for Tesseract OCR fallback.
What: Open-source OCR engine by Google. Reads text from images. Supports eng+msa+chi_sim.
Why: Handles scanned/image-based PDF pages that pypdf cannot extract text from.
What: Python port of Google's language ID library. Probabilistic detection from short text samples.
Why: Fallback language detection when Malay/CJK keyword matching is inconclusive.
What: Commercial TTS API with multilingual v2 model. High-quality, natural-sounding speech.
Why: Reads answers back to users in Voice I/O mode. Falls back to browser speechSynthesis on quota exceed.
What: OpenAI Whisper model hosted on Groq. Transcribes audio to text at high speed.
Why: Converts browser MediaRecorder WebM/Opus audio into question text for the voice input feature.
What: Per-IP rate limiting middleware for FastAPI, built on limits + Redis-compatible backends.
Why: Prevents API abuse. BOOTH_MODE=true loosens limits for demo events where all visitors share one IP.
Frontend & Deployment
What: React framework with App Router, server components, file-based routing, SSR/SSG.
Why: Full frontend framework. Routes map to /, /workspace, /manage, /eval, /about.
What: Typed JavaScript superset. Catches errors at compile time rather than runtime.
Why: Type safety across all components and API response shapes.
What: Utility-first CSS framework. Style via class names, no separate CSS files.
Why: All styling. oklch(0.38 0.13 145) civic green as primary color.
What: Copy-paste components built on Radix UI primitives. Accessible, unstyled base.
Why: Buttons, cards, dialogs, sliders, tabs — accessible foundations with full style control.
What: React animation library with declarative motion primitives and gesture support.
Why: Scroll-triggered fade-ins, hero parallax, card hover effects on landing page.
What: Next.js hosting platform with global edge CDN and automatic deployments from git.
Why: Frontend served globally. Auto-deploys on push to master.
What: Cloud PaaS for containerised apps. Free tier with 512MB RAM.
Why: Hosts the FastAPI backend. requirements.txt strips torch/transformers to fit RAM limits.
Key Features
- Cohere embed-multilingual-v3.0 handles 100+ languages in one vector space
- Language auto-detected per question — answer in same language
- Query augmented into all 3 supported languages simultaneously
- No re-ingestion needed when switching languages
- 3-tier confidence system: strong / cautious / refuse
- Never generates answers from outside the uploaded document
- Faithfulness score measures how grounded the answer is
- Every citation shows page number, section, and scores
- STT: MediaRecorder → Groq Whisper → transcribed question
- TTS: answer → ElevenLabs multilingual v2 → MP3
- Graceful fallback to browser native speechSynthesis
- Language reservation param ready for per-language voice
- ROUGE-1/2/L, BLEU, FK grade — all computed in-house
- Semantic similarity via Cohere embeddings
- Streaming test suite — results appear question-by-question
- Latency tracking: p50/p95/p99 per language
- Upload: validates PDF → ingests → stores in Pinecone + Supabase
- Delete: removes vectors from Pinecone namespace + file from Supabase
- Rename: updates doc_name in all Pinecone vector metadata chunks
- LRU cache (200-entry) invalidated per-document on delete/rename
Evaluation Metrics
All metrics computed in-house in utils/evaluation.py. No external eval APIs.
Recall-Oriented Understudy for Gisting Evaluation. Measures n-gram overlap between generated and reference answers.
ROUGE-1 = unigrams, ROUGE-2 = bigrams, ROUGE-L = longest common subsequence. F1 score reported.
Bilingual Evaluation Understudy. Standard machine translation metric measuring precision of n-gram matches.
Modified n-gram precision with brevity penalty. Originally from MT, adapted here for generation quality.
How grounded the generated answer is in the retrieved source chunks.
Computed by passing the answer as a query back through Cohere reranker against the source chunks. Max relevance_score = faithfulness.
Composite score from vector similarity + Cohere rerank score for the top retrieved chunk.
final = vector_score × 0.35 + rerank_score × 0.65. Thresholds: ≥0.50 = strong, 0.12–0.49 = cautious, <0.12 = refuse.
Cosine similarity between generated answer and ground truth using Cohere embeddings.
Both texts embedded with embed-multilingual-v3.0 (clustering mode). Dot product / (|a| × |b|).
Readability score estimating US school grade level required to understand the answer.
Lower = simpler language. Target: ≤ grade 8 for accessible civic communication.
End-to-end response time from question receipt to answer complete, in milliseconds.
Percentile breakdown: p50 = median, p95 = 95th percentile, p99 = tail latency. Tracked per question.
github.com/apiz23/lingua-rakyat