On-Device RAG Benchmark
Developed for
Privacy-sensitive products, offline field tools, and teams shipping Arabic assistants on-device
The challenge
- Small on-device models fail open-book QA without retrieval scaffolding
- Cloud RAG breaks offline, latency, and data-residency requirements
- Few public Arabic/edge benchmarks show how far weak LLMs can go with good RAG
The Solution
A modular on-device RAG path—chunking, retrieval, and prompting—benchmarked against device manuals so weak LLMs reach near-ceiling accuracy when grounded in local context.
Key capabilities include:
- ◇ Edge-friendly retrieval and context packing
- ◇ Benchmarks showing weak LLMs approaching 100% on grounded QA
- ◇ Patterns that pair with distillation and SLLM deployments
Impact
Proves that privacy-first Arabic assistants can be accurate on-device when retrieval is first-class.
Organization benefits:
- Keep sensitive documents on the device or premise
- Cut cloud cost and latency for field assistants
- Clear eval story for product and procurement teams
Tags
Research Team
Meet Our PIs
Discover the principal investigator behind this project and the expertise that made it possible.
Hamza Salem
Head of PYXON Labs
Leads PYXON Labs research across Arabic AI, edge systems, governance, and applied products that ship into real environments.
// Open Vacancies
Join as a Scientist
Join our team working on On-Device RAG Benchmark. Explore opportunities in machine learning, Arabic NLP, computer vision, edge systems, and applied AI research.
View Open PositionsApply to this project
Submit the same scientist application used on PYXON Labs — tell us about your CV and how you’d contribute to On-Device RAG Benchmark.
Or email info@pyxon.com
Related Projects
Explore more PYXON Labs research connected to this work.