Leaderboard

Eighteen models, open-weight and closed-source, on 879 Bengali harm prompts. Lower is safer. Every model is more likely to answer when the same request is written as formal Bengali journalism than as a casual message, which is the column on the far right.

Columns

Scroll the table sideways for the remaining columns. Rank and model stay put.

What each column means

Add your model

The CLI runs the whole benchmark against any OpenAI-compatible endpoint and prints the same numbers you see here, positioned against this cohort. Email the generated report.json to naymul504@gmail.com or sroydip1@umbc.edu and we will add the row.

uvx banglasafe run \
  --model your-model \
  --base-url http://localhost:8000/v1 \
  --judge-model anthropic/claude-opus-4-7 \
  --judge-base-url https://openrouter.ai/api/v1