# Arabic AI Atlas > The Arabic AI ecosystem as a map, a list, and a skill your agent can install. 3302 entries, generated 2026-10-05. ## Models - [XTTS-v2](https://huggingface.co/coqui/XTTS-v2): Coqui's multilingual voice cloning, Arabic support, 6s cloning - [whisper-large-v3-turbo](https://huggingface.co/openai/whisper-large-v3-turbo): 4x faster Whisper, Arabic support, MIT license - [openai/whisper-large-v3](https://huggingface.co/openai/whisper-large-v3): Supports Arabic among many languages - [Multilingual Chatterbox](https://huggingface.co/ResembleAI/chatterbox): Resemble AI Chatterbox multilingual TTS model covering 22 languages including Arabic. - [wav2vec2-large-xlsr-53-arabic](https://huggingface.co/jonatasgrosman/wav2vec2-large-xlsr-53-arabic): Fine-tuned on Common Voice & Arabic Speech Corpus - [AraBERTv02](https://huggingface.co/aubmindlab/bert-base-arabertv02): Improved tokenization (135M params, 12M+ downloads) - [Phi-4](https://huggingface.co/microsoft/phi-4): Multilingual with Arabic, compact & efficient - [gemma4 e4b claims comparison](https://huggingface.co/k-chirkunov/gemma4-e4b-claims-comparison): Gemma 4 E4B LoRA fine-tune for Arabic dialectal claims comparison (NLI). - [Llama 3.3](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct): Strong Arabic performance - [MMS-1b-all](https://huggingface.co/facebook/mms-1b-all): Meta's Massively Multilingual Speech, ASR for 1100+ languages - [SeamlessM4T v2](https://huggingface.co/facebook/seamless-m4t-v2-large): Meta's all-in-one ASR + translation, ~100 languages inc. Arabic - [Fanar-2-27B-Instruct](https://huggingface.co/QCRI/Fanar-2-27B-Instruct): Fanar 2 Arabic-English instruct model built on Gemma 3 27B, released March 2026. - [Voxtral Mini](https://huggingface.co/mistralai/Voxtral-Mini-3B-2507): Mistral's speech model, 3B, Arabic support, Apache 2.0 - [opus-mt-ar-en](https://huggingface.co/Helsinki-NLP/opus-mt-ar-en): Translation AR→EN - Helsinki-NLP (12.4M+ downloads) - [AraPoemBERT](https://huggingface.co/faisalq/bert-base-arapoembert): BERT model pretrained on Arabic poetry for meter, rhyme, sentiment classification. - [bert-base-arabic-camelbert-mix-sentiment](https://huggingface.co/CAMeL-Lab/bert-base-arabic-camelbert-mix-sentiment): CAMeLBERT-Mix fine-tuned for Arabic sentiment analysis. - [MARBERTv2](https://huggingface.co/UBC-NLP/MARBERTv2): Updated with improved dialectal coverage - [SentimentArEng](https://huggingface.co/qandos0/SentimentArEng): Fine-tune of cardiffnlp twitter-xlm-roberta-base-sentiment for Arabic-English sentiment classification. - [bert-base-arabic-camelbert-da-sentiment](https://huggingface.co/CAMeL-Lab/bert-base-arabic-camelbert-da-sentiment): CAMeLBERT-DA fine-tuned for Arabic sentiment analysis. - [cohere-transcribe-arabic-07-2026](https://huggingface.co/CohereLabs/cohere-transcribe-arabic-07-2026): Cohere Transcribe specialised for Arabic speech recognition. - [opus mt tc big ar en](https://huggingface.co/Helsinki-NLP/opus-mt-tc-big-ar-en): Neural machine translation model for translating from Arabic (ar) to English (en). - [opus-mt-en-ar](https://huggingface.co/Helsinki-NLP/opus-mt-en-ar): Translation EN→AR - Helsinki-NLP (3.5M+ downloads) - [whisper-large-v3-turbo-ar-quran](https://huggingface.co/naazimsnh02/whisper-large-v3-turbo-ar-quran): Whisper large-v3-turbo fine-tuned on Quran recitation. - [qwen3-asr-arabic-uae](https://huggingface.co/vadimbelsky/qwen3-asr-arabic-uae): Qwen3-ASR fine-tuned for Emirati Arabic. - [Whisper Quran](https://huggingface.co/tarteel-ai/whisper-base-ar-quran): Whisper fine-tuned for Quranic recitation recognition ## Datasets - [prophet mosque library](https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library): Prophet’s Mosque Library is one of the primary resources for Islamic books. - [Waqfeya Library](https://huggingface.co/datasets/ieasybooks-org/waqfeya-library): Book texts and PDFs from the Waqfeya Islamic library, 10K-100K files. - [shamela waqfeya library](https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library): Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. - [SARD](https://huggingface.co/datasets/riotu-lab/SARD): Synthetic Arabic OCR dataset for recognition training - [Arabic Books](https://huggingface.co/datasets/MohamedRashad/arabic-books): 8,500 rows of full Arabic book texts. - [XNLI](https://huggingface.co/datasets/facebook/xnli): Evaluation set for XLU by extending the development and test sets of the Multi-Genre Natural Language Inference Corpus (MultiNLI) to 15 languages, - [Shamela4 Full DB](https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB): A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. - [wikiann](https://huggingface.co/datasets/unimelb-nlp/wikiann): Name tagging and linking for 282 languages from Wikipedia, with an Arabic subset. - [Aya Dataset](https://huggingface.co/datasets/CohereForAI/aya_dataset): The Aya Dataset is a multilingual instruction fine-tuning dataset curated by an open-science community via Aya Annotation Platform from Cohere For AI. - [Global-MMLU Lite](https://huggingface.co/datasets/CohereLabs/Global-MMLU-Lite): Lite version of Global-MMLU with 200 culturally sensitive and 200 culturally agnostic samples per language across 16 languages, including Arabic. - [Quranic Recitation Data](https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data): Large Quran recitation audio dataset with 100K-1M clips. - [Arabic Common Voice](https://huggingface.co/datasets/legacy-datasets/common_voice): An open source, multi-language dataset of voices that anyone can use to train speech-enabled applications. - [coda llm data](https://huggingface.co/datasets/mohameddalii/coda-llm-data): This repository contains the full end-to-end dataset, fine-tuning scripts, evaluation suites, load testing harness, and proxy architecture for Coda LLM. - [MMMLU](https://huggingface.co/datasets/openai/MMMLU): Translated the MMLU’s test set into 14 languages using professional human translators. - [XStoryCloze](https://huggingface.co/datasets/juletxara/xstory_cloze): XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages. - [MASSIVE](https://huggingface.co/datasets/qanastek/MASSIVE): 1M parallel labeled virtual-assistant utterances in 51 languages, with Arabic as a subset. - [TYDIQA](https://huggingface.co/datasets/google-research-datasets/tydiqa): Question answering dataset covering 11 typologically diverse languages with 200K question-answer pairs - [Universal Dependencies](https://huggingface.co/datasets/universal-dependencies/universal_dependencies): Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages. - [quranic universal ayahs](https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs): Qur'anic ayah recitation audio with forced-alignment word and letter timestamps for ASR and segmentation. - [Dialectal Arabic Lahgtna v2](https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2): Multi-dialect Arabic speech across 13 dialects, 3,000+ hours. - [Quranic Translation Audio Data](https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data): Complete Quran translation audio and commentary recitations across 51 translation directories in 14 languages. - [prophet mosque library compressed](https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed): Prophet’s Mosque Library is one of the primary resources for Islamic books. - [Quranic Word By Word Audio Data](https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data): Two word-by-word Quran recitation sets, Muallim (teacher style) and Mujawwad (tajweed style), for ASR and TTS. - [waqfeya library compressed](https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed): Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. - [arwiki](https://huggingface.co/datasets/CALM/arwiki): This dataset is extracted using wikiextractor tool, from Wikipedia Arabic pages. ## Tools & Benchmarks - [Global-MMLU](https://huggingface.co/datasets/CohereForAI/Global-MMLU): Multilingual MMLU-style benchmark of 42 languages (Arabic included) with culturally sensitive and culturally agnostic subsets. - [FLORES-101](https://huggingface.co/datasets/gsarti/flores_101): Low-resource machine translation benchmark of 101 languages, with Arabic as one of them. - [xquad](https://huggingface.co/datasets/google/xquad): Cross-lingual QA benchmark of 240 SQuAD paragraphs and 1190 question-answer pairs translated into ten languages including Arabic. - [TunisianMMLU](https://huggingface.co/datasets/linagora/TunisianMMLU): MMLU translated into Tunisian Derja, LINAGORA (France); evaluated with lighteval. - [ArabicMMLU](https://huggingface.co/datasets/MBZUAI/ArabicMMLU): Multi-task language understanding from school exams - [Belebele-Fleurs](https://huggingface.co/datasets/WueNLP/belebele-fleurs): Belebele-Fleurs extends BeleBele into a spoken SLU benchmark with multilingual multiple-choice listening comprehension QA from speech. - [DarijaMMLU](https://huggingface.co/datasets/MBZUAI-Paris/DarijaMMLU): MMLU translated into Moroccan Darija, 22k+ multiple-choice questions. - [AlGhafa Arabic LLM Benchmark (Native)](https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native): Native-Arabic multiple-choice evaluation tasks used as OALL leaderboard tasks. - [OALL Arabic MMLU](https://huggingface.co/datasets/OALL/Arabic_MMLU): GPT-translated Arabic MMLU copy from FreedomIntelligence, used as an OALL v1 task. - [Human-Translated Arabic MMLU](https://huggingface.co/datasets/MBZUAI/human_translated_arabic_mmlu): Human-translated MMLU into Arabic from MBZUAI, companion to ArabicMMLU. - [MIRACL-VISION](https://huggingface.co/datasets/nvidia/miracl-vision): Multilingual visual document retrieval benchmark extending MIRACL with Wikipedia page images for 18 languages. - [Arabic EXAMS (OALL)](https://huggingface.co/datasets/OALL/Arabic_EXAMS): Arabic subset of the EXAMS multilingual school-exam benchmark, an OALL task. - [EgyMMLU](https://huggingface.co/datasets/UBC-NLP/EgyMMLU): MMLU translated into Egyptian Arabic from UBC (Canada), released with NileChat. - [ASR Code Switch](https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch): A curated benchmark of 1,200 code-switching utterances (300 per language pair) for evaluating commercial ASR systems on multilingual speech. - [artelingo dummy](https://huggingface.co/datasets/youssef101/artelingo-dummy): ArtELingo is a benchmark and dataset introduced in a research paper aimed at promoting work on diversity across languages and cultures. - [Arabic MMLU 10percent](https://huggingface.co/datasets/arcee-globe/Arabic_MMLU-10percent): 10 percent subset of Arabic MMLU multiple-choice questions across subjects. - [ArAD](https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/ArAD): Benchmark-ready packaging of the test split of the Arabic Audio Deepfake (ArAD) dataset: binary anti-spoofing on Arabic (primarily Levantine dialect) speech. - [Quranic ASR Benchmark](https://huggingface.co/datasets/Quran-Lab/quranic-asr-benchmark): Small benchmark set for evaluating ASR models on Quran recitation. - [MMCQA-SemEval27](https://huggingface.co/datasets/QCRI/MMCQA-SemEval27): This is the dataset for MMCultureQA, the SemEval 2027 shared task on culturally grounded visual question answering in English and Arabic. - [ALRAGE](https://huggingface.co/datasets/OALL/ALRAGE): Arabic retrieval-augmented generation evaluation set used in OALL v2. - [AraTrust](https://huggingface.co/datasets/asas-ai/AraTrust): Arabic LLM trustworthiness benchmark across truthfulness, ethics, safety and privacy. - [AraDICE ArabicMMLU Egyptian](https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy): Egyptian-dialect translation of ArabicMMLU from the QCRI AraDiCE benchmark suite. - [jev ar bench](https://huggingface.co/datasets/atmaneayoub/jev-ar-bench): Arabic intent-routing benchmarks for MSA, Emirati, Saudi and code-switched Arabic - [fa en ar handwritten ocr v1](https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1): A large, clean, augmentation-rich synthetic handwriting dataset for training and benchmarking OCR / HTR models on Persian (fa), Arabic (ar) and English (en). - [ALM-Bench](https://huggingface.co/datasets/MBZUAI/ALM-Bench): All Languages Matter multimodal cultural benchmark covering 100 languages including Arabic. ## Papers - [Systematic Review and Meta-Analysis of 2024–2025 Studies on AI Arabic Translation, Linguistics and Pedagogy](https://doi.org/10.32996/jcsts.2026.5.1.2): This study aims to conduct a systematic review (SR) and meta-analysis (MA) of twenty articles by the author published between 2024–2025 on the use of AI models - [A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model](https://arxiv.org/abs/2610.00223): As ever larger language models grow more ubiquitous, it is crucial to consider their environmental impact. - [A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition](https://arxiv.org/abs/2606.19747): This paper presents a systematic empirical study of domain-specific fine-tuning of pretrained Transformer-based models for Quranic ASR. - [A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator](https://arxiv.org/abs/2609.14967): On top of the normalized text, we build a deterministic, LLM-free Quranic recitation validator using a four-layer verse-matching search. - [A Human-in-the-Loop Label Error Detection Framework Applied to Arabic-Script HTR Datasets](https://arxiv.org/abs/2601.16713): To help closing this gap, we propose a two-stage framework (CER-HV) for detecting label errors. - [Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education](https://arxiv.org/abs/2603.20255): However, children speech research remains limited due to the lack of publicly available datasets. - [ADAB: Arabic Dataset for Automated Politeness Benchmarking -- A Large-Scale Resource for Computational Sociopragmatics](https://arxiv.org/abs/2602.13870): We introduce ADAB (Arabic Politeness Dataset), a new annotated Arabic dataset collected from four online platforms, including social media, e-commerce. - [AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes](https://arxiv.org/abs/2607.27393): We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained. - [Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs](https://github.com/UBC-NLP/Alexandria): We introduce Alexandria, a large-scale, community-driven, human-translated dataset designed to bridge this gap. - [AlexandriaX 2026: The First Shared Task on Dialectal Arabic Machine Translation](https://arxiv.org/abs/2609.22796): We present the AlexandriaX 2026 Shared Task on Dialectal Arabic MT, which addresses these challenges through three complementary subtasks. - [Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models](https://github.com/qcri/Almieyar-Oryx-BloomBench): We introduce BloomBench, part of the Almieyar benchmarking series, the first cognitively human-grounded, bilingual. - [Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition](https://arxiv.org/abs/2609.35564): We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families. - [An End-to-End Hybrid Framework for Rumour Detection in Low-Resources Algerian Dialect](https://arxiv.org/abs/2606.13411): This paper presents an end-to-end rumour detection hybrid framework for Algerian dialect social media content. - [ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination](https://arxiv.org/abs/2605.22081): We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination. - [Arabic Prompts with English Tools: A Benchmark](https://arxiv.org/abs/2601.05101): Large Language Models (LLMs) are now integral to numerous industries, increasingly serving as the core reasoning engine for autonomous agents. - [ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform](https://arxiv.org/abs/2601.22987): We present ArabicDialectHub, a cross-dialectal Arabic learning resource comprising 552 phrases across six varieties. - [ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification](https://arxiv.org/abs/2608.01291): ArabicDialectSafety: 25,071 human-curated safety prompts across MSA and four dialect groups, with seven harm categories and a dual-task evaluation. - [AraDetox: A Multi-Dialect Arabic Detoxification Dataset](https://github.com/ArabicNLP-UK/AraDetox): We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified rewrites. - [ARAFA: An LLM-Generated Arabic Fact-Checking Dataset](https://arxiv.org/abs/2609.25833): We introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework. - [AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task](https://arxiv.org/abs/2609.27387): AraGenre is a shared task on hierarchical, definition-guided Arabic genre classification, motivated by the limited availability of annotated data in Arabic. - [AraHopeCorpus: Annotation Guidelines and Dataset for Hope Speech in Arabic Social Media Crisis Discourse](https://arxiv.org/abs/2605.23325): This paper introduces AraHopeCorpus, the first annotated dataset of Arabic hope speech collected from ten thousand YouTube comments related to the war. - [AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations](https://arxiv.org/abs/2608.26921): We introduce AraMS-28k, the largest publicly released line-level dataset of genuine historical Arabic manuscripts, comprising 14 books, 3,043 pages. - [ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging](https://huggingface.co/datasets/riotu-lab/ARCADE-full): We present ARCADE (Arabic Radio Corpus for Audio Dialect Evaluation), the first Arabic speech dataset designed explicitly with city-level dialect granularity. - [Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation](https://arxiv.org/abs/2604.03395): We present QIMMA, a quality-assured Arabic LLM leaderboard that places systematic benchmark validation at its core. - [ArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization](https://arxiv.org/abs/2605.20967): This paper presents ArPoMeme, a large-scale dataset of approximately 7,300 Arabic political memes categorized by ideological orientation. ## Organizations - [King Saud University](https://huggingface.co/faisalq/SaudiBERT): SaudiBERT, Saudi dialect corpora (STMC, SFC) - [AIDA Lab (PSU)](https://aidalabpsu.com/): AI and Data Analytics lab at Prince Sultan University, Riyadh. - [AIM Lab (NU-Q)](https://www.qatar.northwestern.edu/research/aim-lab/): AI and media research lab at Northwestern University in Qatar. - [Ain Shams University](https://cis.asu.edu.eg/): Arabic NLP, sentiment analysis, NER research - [AlooChat](https://aloochat.ai): AI customer support chat agents for WhatsApp, Instagram, web and email, found via an Arabic conversational AI search in Oman. - [AlQari](https://alqari.sa): Arabic-first document intelligence turning Arabic and English documents into structured, validated data. - [Applied Innovation Center (AIC)](https://aic.gov.eg): Egyptian MCIT AI center; builds the Karnak LLM family and publishes models on Hugging Face (Applied-Innovation-Center). - [Arabic.AI (Tarjama)](https://www.tarjama.com/): Arabic-first autonomous AI - Pronoia Arabic LLM, Agentic AI platform - [Arabot](https://arabot.io/): Conversational AI for Arabic - Arabic NLP chatbot engine - [Aramco Digital](https://www.aramcodigital.com): Aramco digital arm, behind Metabrain generative AI assistant and AI partnerships. - [ARBML](https://github.com/ARBML): Democratizing Arabic NLP - masader, klaam, tkseem - [ASAS AI](https://asas.ai): Saudi AI company offering Arabic models, automation and in-Kingdom infrastructure. - [Astra Tech (Botim)](https://astratech.ae): Abu Dhabi tech group behind Botim, building AI assistant products. - [AtlasIA](https://www.atlasia.ma/): Behind AL Atlas Moroccan Darija models - [AUB MIND Lab](https://github.com/aub-mind): Foundational Arabic NLP models - AraBERT, AraGPT2, AraELECTRA - [Calfa](https://calfa.fr): OCR service extracting printed and handwritten text from scanned documents, including Arabic script. - [CAMeL Lab (NYU Abu Dhabi)](https://camel-lab.com/): CAMeLBERT, camel_tools, morphological analysis - [CAMeL Lab, NYUAD](https://github.com/CAMeL-Lab): Arabic NLP tools and models - CAMeLBERT, camel_tools - [Clusterlab AI](https://huggingface.co/ClusterlabAi): Behind 101 Billion Arabic Words Dataset and InstAr-500k - [Cohere](https://cohere.com/): Multilingual LLMs - Command R Arabic, RAG optimization - [Convertedin](https://www.converted.in/): AI marketing automation - Arabic/English e-commerce personalization, $3M funded - [Core42](https://www.core42.ai): G42 sovereign cloud and AI company; Compass AI platform and Jais deployment. - [Crowd Analyzer](https://crowdanalyzer.com/): Arabic social media monitoring - Arabic NLP analytics, sentiment analysis, media monitoring - [Dar Al-Raqmana](https://daralraqmana.com): Digitization company offering Arabic text recognition (OCR), digital repositories and morphological search. - [Darijat TTS](https://tts.darijat.com/): Arabic dialect TTS service claiming 23 dialects. ## Agent Skills - [ArabAgentSkills](https://github.com/ArabAgentSkills/Skills): Public source-backed Arab-world agent skills covering APIs, marketing, sales, support and localization. - [ArabGuard](https://github.com/d12o6aa/arabguard): Python SDK protecting LLMs and chatbots from prompt injection in Arabic text. - [Arabic Content Studio](https://github.com/smeseik-ai/arabic-content-studio): Claude Skills for Arabic content teams covering production, culture, SEO and brand voice. - [Arabic Dict MCP](https://github.com/arnizamani/arabic-dict-mcp): MCP server for grounded Arabic dictionary lookup wrapping arramooz (MSA verbs and nouns). - [Arabic Scholar MCP Server](https://github.com/EngDawood/arabic-scholar-mcp-server): MCP server for searching Arabic academic research, articles and dissertations across sources. - [Arabic Video Subtitles Skill](https://github.com/EngDawood/arabic-video-subtitles-skill): Skill with Netflix-style Arabic subtitling guidelines and an SRT/VTT checker. - [Arabic Word Production Skill](https://github.com/Bannovich/arabic-word-production): Deterministic agent skill and plugin for Arabic-first and bilingual Word DOCX production. - [arabic-bidi-engineering](https://github.com/MosaabGalmod/arabic-bidi-engineering): Agent skill for correct Arabic RTL/BiDi in chat, terminal and documents - [arabic-pii-py](https://github.com/Aajil-Labs/arabic-pii-py): Local-first reversible PII tokenization for Arabic and Gulf data before it reaches an LLM. - [ArabiMaak MCP](https://github.com/deepdiver4ai/arabimaak-mcp): Gulf Arabic language MCP server for word, phrase and cultural-context lookup. - [awesome-arabic-claude-skills](https://github.com/EngDawood/awesome-arabic-claude-skills): Curated library of open-source Arabic skills for Claude Code and agents - [Azan MCP](https://github.com/ahmedeltaher/azan-mcp): Lightweight MCP library for Islamic prayer times and Qibla for AI agents; not Arabic-language specific. - [Claude Adhkar](https://github.com/nosseralaa7-rgb/claude-adhkar): Shows Arabic-script adhkar in the Claude Code spinner with a terminal font that renders Arabic. - [Claude Arabic Writing](https://github.com/ahmeddabak/claude-arabic-writing): Claude Agent Skill for natural, grammatically correct Arabic writing and translation. - [Claude Code RTL Extension](https://github.com/yechielby/claude-code-rtl-extension): VS Code and Cursor extension adding RTL support for Hebrew and Arabic in Claude Code. - [Claude Desktop RTL Patch (macOS)](https://github.com/toboly/claude-desktop-rtl-patch-mac): Adds auto-detected Hebrew and Arabic RTL support to Claude Desktop on macOS. - [Dorar Hadith MCP](https://github.com/ibnsaleem29/dorar-hadith-mcp): Claude extension and MCP server for Hadith research including isnad and takhrij via Dorar. - [Fanar MCP Server](https://github.com/danijeun/fanar-mcp-server): MCP server exposing Fanar API tools such as Islamic RAG and image generation. - [Hadith MCP (ovehbe)](https://github.com/ovehbe/hadith-mcp): MCP server for searchable, citation-safe hadith text. - [Hurmoz](https://github.com/Moshe-ship/hurmoz): Collection of 63 Arabic skills for the Hermes Agent framework. - [Islamic Knowledge Skill](https://github.com/shadysalman/islamic-knowledge-skill): RAG-powered Claude skill over the Quran (6,236 ayat) and Sahih Bukhari. - [karem-arabic-presentation](https://github.com/karem505/karem-arabic-presentation): Claude Code skill for Arabic/English bilingual RTL HTML presentations - [Kazma](https://github.com/Mubder/kazma): Self-hosted agent platform, bilingual English and Arabic by design. - [Kivun Terminal WSL](https://github.com/noambrand/kivun-terminal-wsl): Claude Code in WSL with Hebrew, Arabic and Persian rendered correctly. - [MasrKit](https://github.com/asasemahmed/MasrKit): Open skills for coding agents to build products that feel Egyptian. ## Optional - [Full dataset (JSON)](https://raw.githubusercontent.com/h9-tec/arabic-ai-atlas/main/dist/atlas.json): every entry with metadata and metrics - [README](https://github.com/h9-tec/arabic-ai-atlas#readme): the human-readable atlas with every table