AtlasIA

community

https://atlasia.ma/

atlasia-ma

Activity Feed

AI & ML interests

Open-source data and models for Morocco.

Recent Activity

BounharAbdelaziz updated a dataset about 15 hours ago

atlasia/FineWeb2-Moroccan-Arabic-Predictions-0.6

BounharAbdelaziz updated a dataset about 15 hours ago

atlasia/Terjman_v2_on_TerjamaBench

BounharAbdelaziz updated a dataset about 16 hours ago

atlasia/FineWeb2-Moroccan-Arabic-Predictions-0.7

View all activity

atlasia's activity

BounharAbdelaziz

updated 2 datasets about 15 hours ago

atlasia/FineWeb2-Moroccan-Arabic-Predictions-0.6

Viewer • Updated about 15 hours ago • 28.5M

atlasia/Terjman_v2_on_TerjamaBench

Viewer • Updated about 15 hours ago • 850

BounharAbdelaziz

updated 3 datasets about 16 hours ago

updated a dataset 1 day ago

atlasia/FineWeb2-Moroccan-Arabic-Predictions-model_binary_v3_1fpr.bin

Viewer • Updated 1 day ago • 69.6M • 3

BounharAbdelaziz

updated a Space 3 days ago

Running

🎙️

Moroccan Fast Speech To Text Transcription

Speech-to-Text Transcription for the moroccan darija dialect

nouamanetazi

updated a dataset 4 days ago

atlasia/chatbot-arena-db

Updated 4 days ago • 4

alielfilali01

posted an update 5 days ago

Post

1673

~75% on the challenging GPQA with only 40M parameters 🔥🥳

GREAT ACHIEVEMENT ! Or is it ?

This new Work, "Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation", take out the mystery about many models i personally suspected their results. Speacially on leaderboards other than the english one, Like the Open Arabic LLM Leaderbaord OALL/Open-Arabic-LLM-Leaderboard.

The authors of this work, first started by training a model on the GPQA data, which, unsurprisingly, led to the model achieving 100% performance.

Afterward, they trained what they referred to as a 'legitimate' model on legitimate data (MedMCQA). However, they introduced a distillation loss from the earlier, 'cheated' model.

What they discovered was fascinating: the knowledge of GPQA leaked through this distillation loss, even though the legitimate model was never explicitly trained on GPQA during this stage.

This raises important questions about the careful use of distillation in model training, especially when the training data is opaque. As they demonstrated, it’s apparently possible to (intentionally or unintentionally) leak test data through this method.

Find out more: Data Laundering: Artificially Boosting Benchmark Results through Knowledge Distillation (2412.15255)

1 reply

nouamanetazi

updated a Space 6 days ago

Running

🚀

darija-chatbot-arena

BounharAbdelaziz

updated a Space 7 days ago

Running

👁

Open Arabic Dialect Identification Leaderboard

BounharAbdelaziz

updated a dataset 8 days ago

atlasia/Arabic-LID-Leaderboard

Viewer • Updated 8 days ago • 234k • 38

alielfilali01

posted an update 22 days ago

Post

3383

Unpopular opinion: Open Source takes courage to do !

Not everyone is brave enough to release what they have done (the way they've done it) to the wild to be judged !
It really requires a high level of "knowing wth are you doing" ! It's kind of a super power !

Cheers to the heroes here who see this!

3 replies

alielfilali01

posted an update 26 days ago

Post

1505

Apparently i forgot to put this here !

Well, this is a bit late but consider given our recent blog a read if you are interested in Evaluation.

You don't have to be into Arabic NLP in order to read it, the main contribution we are introducing is a new evaluation measure for NLG. We made the fisrt application of this measure on Arabic for now and we will be working with colleagues from the community to expand it to other languages.

Blog:
Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard
https://huggingface.co/blog/leaderboard-3c3h-aragen

Space:
inceptionai/AraGen-Leaderboard

Give it a read and let me know your thoughts 🤗

alielfilali01

posted an update about 2 months ago

Post

2185

Unpopular opinion : o1-preview is more stupid than 4o and Qwen2.5-72B-Instruct in extremely underrated !

2 replies

alielfilali01

posted an update 2 months ago

Post

1702

I feel like this incredible resource hasn't gotten the attention it deserves in the community!

@clefourrier and generally the HuggingFace evaluation team put together a fantastic guidebook covering a lot about 𝗘𝗩𝗔𝗟𝗨𝗔𝗧𝗜𝗢𝗡 from basics to advanced tips.

link : https://github.com/huggingface/evaluation-guidebook

I haven’t finished it yet, but i'am enjoying every piece of it so far. Huge thanks @clefourrier and the team for this invaluable resource !

3 replies

alielfilali01

posted an update 3 months ago

Post

1829

Why nobdoy is talking about the new training corpus released by MBZUAI today.

TxT360 is +15 Trillion tokens corpus outperforming FineWeb on several metrics. Ablation studies were done up to 1T tokens.

Read blog here : LLM360/TxT360
Dataset : LLM360/TxT360

2 replies

alielfilali01

posted an update 3 months ago

Post

2571

Don't you think we should add a tag "Evaluation" for datasets that are meant to be benchmarks and not for training ?

At least, when someone is collecting a group of datasets from an organization or let's say the whole hub can filter based on that tag and avoid somehow contaminating their "training" data.

imomayiz

authored a paper 3 months ago

Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dialect

Paper • 2409.17912 • Published Sep 26, 2024 • 23

alielfilali01

posted an update 3 months ago

Post

873

We need a fork feature for models and datasets similar to "Duplicate this space" in spaces ! Don't you think ?

Sometimes you just want to save something in your profile privately and work on it later without the hassle of "load_.../push_to_hub" in a code file.

I know this is super lazy 😅 But it is what it is ...

tag : @victor

5 replies

AI & ML interests

Recent Activity

Team members 14

atlasia's activity

Moroccan Fast Speech To Text Transcription

darija-chatbot-arena

Open Arabic Dialect Identification Leaderboard