Arabic Poetry Benchmark
Developed for
Researchers and product teams building culturally grounded Arabic generative systems
The challenge
- General LLMs score poorly on classical meter and rhyme constraints
- No shared, product-facing benchmark for specialized Arabic poetry tasks
- Hard to measure progress of domain-specialized models
The Solution
A poetry-focused benchmark covering meter/rhyme tasks that shows general models lagging (around mid-50% ranges in Labs demos) and guides specialized model work.
Key capabilities include:
- ◇ Structured meter and rhyme evaluation tasks
- ◇ Baseline scores for general vs specialized Arabic models
- ◇ Research loop feeding generative and SLLM work
Impact
Turns cultural Arabic generation quality into a measurable R&D target rather than anecdotal demos.
Organization benefits:
- Clear KPIs for specialized Arabic LM research
- Better storytelling for Labs and partners
- Guidance for generative product roadmaps
Tags
Research Team
Meet Our PIs
Discover the principal investigator behind this project and the expertise that made it possible.
Hamza Salem
Head of PYXON Labs
Leads PYXON Labs research across Arabic AI, edge systems, governance, and applied products that ship into real environments.
// Open Vacancies
Join as a Scientist
Join our team working on Arabic Poetry Benchmark. Explore opportunities in machine learning, Arabic NLP, computer vision, edge systems, and applied AI research.
View Open PositionsApply to this project
Submit the same scientist application used on PYXON Labs — tell us about your CV and how you’d contribute to Arabic Poetry Benchmark.
Or email info@pyxon.com
Related Projects
Explore more PYXON Labs research connected to this work.