Technical Reference

v2.0.0

How Lingua Rakyat Works

Architecture reference — ingestion pipeline, Q&A system, tech stack, API, and eval metrics.

Project Overview

The Problem
  • Government PDFs written in legalese — not plain language
  • Long documents with no searchable Q&A interface
  • No equal access across Malay, English, and Chinese speakers
  • Eligibility criteria buried in paragraphs
The Solution
  • Upload any government PDF — queryable instantly
  • Ask in Malay, English, or Mandarin — answer in same language
  • Every answer grounded in document text with page citations
  • Evidence guard refuses to hallucinate — shows confidence

RAG in one sentence: Instead of an LLM guessing from training memory, Lingua Rakyat first searches the actual document for relevant passages, then passes only those passages to the LLM — so the answer is always traceable to a real source.

Ingestion Pipeline

Phase 1

Triggered once when a PDF is uploaded. Runs offline — users can start asking questions as soon as ingestion completes.

1
PDF Validation
  • Checks file type, encryption status, and page count (1–500 pages)
  • Extracts text with pypdf — pages with < 50 chars flagged for OCR
  • Rejects PDFs where > 80% of pages have no extractable text
2
Text Extraction + OCR Fallback
  • pypdf extracts text page-by-page for text-based PDFs
  • PyMuPDF renders low-text pages to bitmap at 2× scale
  • Tesseract OCR reads bitmaps — supports eng + msa + chi_sim
3
Section-Aware Chunking
  • Regex detects section headers: Bahagian, Seksyen, Section, numbered (1.2.3), ALLCAPS
  • Each section becomes its own chunk — preserves document structure
  • Long sections split at 360-word target, 45-word overlap, min 20 words/chunk
4
Cohere Embedding
  • All chunks embedded in batch via Cohere embed-multilingual-v3.0
  • Produces 1024-dim vectors per chunk — language-agnostic representation
  • input_type = search_document (optimised for retrieval, not clustering)
5
Storage
  • Vectors + metadata → Pinecone under document namespace
  • Raw PDF file → Supabase Storage bucket
  • Document record (name, chunk count, timestamps) → Supabase PostgreSQL

Q&A Pipeline

Phase 2

Triggered per question. Streams tokens to the UI as they're generated — first token appears within ~600ms on a warm cache.

1
Language Detection
  • Keyword matching first: Malay words (nak, boleh, saya, mohon) with word-boundary regex
  • CJK character ratio: if > 15% of chars are CJK, classify as zh-cn
  • langdetect library as fallback — maps 20+ dialects to en / ms / zh-cn
2
Multi-Query Augmentation
  • Expands 1 question into up to 4 variants: original + paraphrase + translations
  • Translations generated into all other supported languages
  • Improves retrieval recall — same chunks found regardless of query language
3
Semantic Retrieval
  • All variants embedded in a single Cohere API call (input_type = search_query)
  • Each variant queries Pinecone — top-k chunks per variant
  • Results merged, deduplicated by chunk ID, sorted by composite score
4
Neural Reranking
  • Cohere rerank-multilingual-v3.0 re-scores candidates reading query + document together
  • Final score = vector_score × 0.35 + rerank_score × 0.65
  • Higher precision than cosine similarity alone — cross-encoder architecture
5
Evidence Guard
  • Top chunk score ≥ 0.50 → strong evidence → standard QA prompt
  • Score 0.12–0.49 → cautious mode → answer with caveats
  • Score < 0.12 → hard refusal → canned message, no hallucination
6
Answer Generation
  • Context-aware prompt built in detected language (en/ms/zh-cn)
  • Last 3 conversation turns injected for follow-up awareness
  • Groq LLaMA 3.3 70B streams answer in 3–5 bullet points token-by-token
7
Post-Generation
  • Faithfulness: answer passed back through Cohere reranker vs source chunks
  • LLaMA 3.1 8B generates 3 follow-up question suggestions (non-blocking)
  • Result cached (LRU, 200-entry max) — cache bypassed when chat history present

API Reference

All endpoints rate-limited per IP via SlowAPI. Interactive docs at /docs (Swagger UI).

MethodPath
POST/api/documents/upload
GET/api/documents/
DELETE/api/documents/{id}
POST/api/chat/ask-stream
POST/api/chat/ask
GET/api/chat/history
DELETE/api/chat/history/{doc_id}
POST/api/eval/run-test-suite-stream
GET/api/eval/report
GET/api/eval/data-quality
POST/api/voice/transcribe
POST/api/voice/tts
POST/api/feedback

Tech Stack

Backend

FastAPIAPI Framework

What: Python async web framework. Auto-generates Swagger docs at /docs. High throughput via async I/O.

Why: Serves all endpoints — chat, documents, eval, voice — with automatic OpenAPI documentation for judges.

Groq LLaMA 3.3 70BAnswer Generation

What: Meta's open-source LLM (70B params) running on Groq's custom LPU inference hardware. Extremely fast token generation.

Why: Generates grounded, source-cited answers in Malay, English, or Chinese from retrieved document context.

LLaMA 3.1 8B (Groq)Fast Model

What: Smaller, faster variant of LLaMA for low-latency tasks where full 70B is overkill.

Why: Generates follow-up question suggestions after each answer without blocking the main response stream.

Cohere embed-multilingual-v3.0Embeddings

What: Embedding model supporting 100+ languages. Converts text into 1024-dim vectors capturing semantic meaning.

Why: Embeds both document chunks (at ingestion) and queries (at retrieval) enabling cross-lingual semantic search.

Cohere rerank-multilingual-v3.0Neural Reranking

What: Cross-encoder model that re-scores retrieved candidates by reading query + document together.

Why: Re-ranks vector search results in context — higher precision than pure cosine similarity. Also computes faithfulness.

PineconeVector Database

What: Managed cloud vector database. Stores and searches high-dimensional embeddings at scale.

Why: Stores all document chunk vectors. Each document gets its own namespace for isolated retrieval.

SupabaseStorage & DB

What: Open-source Firebase alternative — PostgreSQL + file storage + real-time subscriptions.

Why: Stores raw PDFs (Storage bucket), document metadata, chat history per user/document, feedback thumbs.

pypdf + PyMuPDFPDF Processing

What: pypdf: pure-Python text extraction. PyMuPDF (fitz): C-based PDF renderer for rasterising pages to images.

Why: pypdf handles text-based PDFs. PyMuPDF renders scanned pages for Tesseract OCR fallback.

Tesseract OCROCR Fallback

What: Open-source OCR engine by Google. Reads text from images. Supports eng+msa+chi_sim.

Why: Handles scanned/image-based PDF pages that pypdf cannot extract text from.

langdetectLanguage Detection

What: Python port of Google's language ID library. Probabilistic detection from short text samples.

Why: Fallback language detection when Malay/CJK keyword matching is inconclusive.

ElevenLabsText-to-Speech

What: Commercial TTS API with multilingual v2 model. High-quality, natural-sounding speech.

Why: Reads answers back to users in Voice I/O mode. Falls back to browser speechSynthesis on quota exceed.

Groq WhisperSpeech-to-Text

What: OpenAI Whisper model hosted on Groq. Transcribes audio to text at high speed.

Why: Converts browser MediaRecorder WebM/Opus audio into question text for the voice input feature.

SlowAPIRate Limiting

What: Per-IP rate limiting middleware for FastAPI, built on limits + Redis-compatible backends.

Why: Prevents API abuse. BOOTH_MODE=true loosens limits for demo events where all visitors share one IP.

Frontend & Deployment

Next.js 15React Framework

What: React framework with App Router, server components, file-based routing, SSR/SSG.

Why: Full frontend framework. Routes map to /, /workspace, /manage, /eval, /about.

TypeScriptType Safety

What: Typed JavaScript superset. Catches errors at compile time rather than runtime.

Why: Type safety across all components and API response shapes.

Tailwind CSSStyling

What: Utility-first CSS framework. Style via class names, no separate CSS files.

Why: All styling. oklch(0.38 0.13 145) civic green as primary color.

shadcn/uiComponent Library

What: Copy-paste components built on Radix UI primitives. Accessible, unstyled base.

Why: Buttons, cards, dialogs, sliders, tabs — accessible foundations with full style control.

Framer MotionAnimations

What: React animation library with declarative motion primitives and gesture support.

Why: Scroll-triggered fade-ins, hero parallax, card hover effects on landing page.

VercelFrontend Deploy

What: Next.js hosting platform with global edge CDN and automatic deployments from git.

Why: Frontend served globally. Auto-deploys on push to master.

RenderBackend Deploy

What: Cloud PaaS for containerised apps. Free tier with 512MB RAM.

Why: Hosts the FastAPI backend. requirements.txt strips torch/transformers to fit RAM limits.

Key Features

Multilingual RAG
  • Cohere embed-multilingual-v3.0 handles 100+ languages in one vector space
  • Language auto-detected per question — answer in same language
  • Query augmented into all 3 supported languages simultaneously
  • No re-ingestion needed when switching languages
Evidence Guard — Anti-Hallucination
  • 3-tier confidence system: strong / cautious / refuse
  • Never generates answers from outside the uploaded document
  • Faithfulness score measures how grounded the answer is
  • Every citation shows page number, section, and scores
Voice I/O
  • STT: MediaRecorder → Groq Whisper → transcribed question
  • TTS: answer → ElevenLabs multilingual v2 → MP3
  • Graceful fallback to browser native speechSynthesis
  • Language reservation param ready for per-language voice
Evaluation Dashboard
  • ROUGE-1/2/L, BLEU, FK grade — all computed in-house
  • Semantic similarity via Cohere embeddings
  • Streaming test suite — results appear question-by-question
  • Latency tracking: p50/p95/p99 per language
Document Management
  • Upload: validates PDF → ingests → stores in Pinecone + Supabase
  • Delete: removes vectors from Pinecone namespace + file from Supabase
  • Rename: updates doc_name in all Pinecone vector metadata chunks
  • LRU cache (200-entry) invalidated per-document on delete/rename

Evaluation Metrics

All metrics computed in-house in utils/evaluation.py. No external eval APIs.

ROUGE-1 / ROUGE-2 / ROUGE-L
0 – 1

Recall-Oriented Understudy for Gisting Evaluation. Measures n-gram overlap between generated and reference answers.

ROUGE-1 = unigrams, ROUGE-2 = bigrams, ROUGE-L = longest common subsequence. F1 score reported.

BLEU
0 – 1

Bilingual Evaluation Understudy. Standard machine translation metric measuring precision of n-gram matches.

Modified n-gram precision with brevity penalty. Originally from MT, adapted here for generation quality.

Faithfulness Score
0 – 1

How grounded the generated answer is in the retrieved source chunks.

Computed by passing the answer as a query back through Cohere reranker against the source chunks. Max relevance_score = faithfulness.

Confidence Score
0 – 1

Composite score from vector similarity + Cohere rerank score for the top retrieved chunk.

final = vector_score × 0.35 + rerank_score × 0.65. Thresholds: ≥0.50 = strong, 0.12–0.49 = cautious, <0.12 = refuse.

Semantic Similarity
–1 – 1

Cosine similarity between generated answer and ground truth using Cohere embeddings.

Both texts embedded with embed-multilingual-v3.0 (clustering mode). Dot product / (|a| × |b|).

Flesch-Kincaid Grade
1 – 18+

Readability score estimating US school grade level required to understand the answer.

Lower = simpler language. Target: ≤ grade 8 for accessible civic communication.

Latency (p50/p95/p99)
ms

End-to-end response time from question receipt to answer complete, in milliseconds.

Percentile breakdown: p50 = median, p95 = 95th percentile, p99 = tail latency. Tracked per question.

Note: This page reflects the live codebase state as of the RISE 2026 submission. Backend hosted on Render (free tier, 512 MB RAM). Frontend on Vercel. Source: github.com/apiz23/lingua-rakyat