← All lab projects

Arabic Poetry Benchmark

Developed for

Researchers and product teams building culturally grounded Arabic generative systems

The challenge

  • General LLMs score poorly on classical meter and rhyme constraints
  • No shared, product-facing benchmark for specialized Arabic poetry tasks
  • Hard to measure progress of domain-specialized models

The Solution

A poetry-focused benchmark covering meter/rhyme tasks that shows general models lagging (around mid-50% ranges in Labs demos) and guides specialized model work.

Key capabilities include:

  • Structured meter and rhyme evaluation tasks
  • Baseline scores for general vs specialized Arabic models
  • Research loop feeding generative and SLLM work

Impact

Turns cultural Arabic generation quality into a measurable R&D target rather than anecdotal demos.

Organization benefits:

  • Clear KPIs for specialized Arabic LM research
  • Better storytelling for Labs and partners
  • Guidance for generative product roadmaps

Tags

BenchmarksPoetryArabic NLPGenerative AI

Research Team

Meet Our PIs

Discover the principal investigator behind this project and the expertise that made it possible.

Hamza Salem

Hamza Salem

Head of PYXON Labs

Leads PYXON Labs research across Arabic AI, edge systems, governance, and applied products that ship into real environments.

// Open Vacancies

Join as a Scientist

Join our team working on Arabic Poetry Benchmark. Explore opportunities in machine learning, Arabic NLP, computer vision, edge systems, and applied AI research.

View Open Positions

Apply to this project

Submit the same scientist application used on PYXON Labs — tell us about your CV and how you’d contribute to Arabic Poetry Benchmark.

Or email info@pyxon.com

Related Projects

Explore more PYXON Labs research connected to this work.