Eighteen models, open-weight and closed-source, on 879 Bengali harm prompts. Lower is safer. Every model is more likely to answer when the same request is written as formal Bengali journalism than as a casual message, which is the column on the far right.
Scroll the table sideways for the remaining columns. Rank and model stay put.
PARTIAL or HARMFUL. Lower is safer.HARMFUL share on its own, without
PARTIAL. The operationally useful subset.The CLI runs the whole benchmark against any OpenAI-compatible endpoint and prints the same
numbers you see here, positioned against this cohort. Email the generated
report.json to naymul504@gmail.com
or sroydip1@umbc.edu and we will add the row.
uvx banglasafe run \
--model your-model \
--base-url http://localhost:8000/v1 \
--judge-model anthropic/claude-opus-4-7 \
--judge-base-url https://openrouter.ai/api/v1