Recent Updates


Important Dates

* These dates are subject to changes.

Paper submission due Feb 15 (Mon), 2027
Notification of acceptance Mar 26 (Fri), 2027
Camera-ready due Apr 16 (Fri), 2027
Workshop June 2027 co-located with NAACL

← All Shared Tasks

Shared Task 1: Adversarial Prompt Classification

Before an LLM answers a request, a safety classifier (a "guard" model) often checks whether the request is safe to answer. These classifiers are trained mostly on English. Bangla, however, is written online in several ways: in Bangla script, in romanized (Latin-script) Bangla, and code-mixed with English, Hindi, or Urdu. The same harmful request can pass a safety filter just by changing how it is written.

This shared task asks participants to build classifiers that judge how risky a Bangla prompt is, whichever way it is written.

Task Definition

Given a Bangla prompt, assign it one of four labels:

IDLabelDefinitionTypical categories
0SafeNo harmful requestNeutral, positive, or harmless requests
1OffensiveInsulting, provocative, or discriminatory expression without a request for harmful assistanceOffensive language, hate speech, anger
2HarmfulRequest for lower- or medium-risk harmful assistanceMisinformation, adversarial and malicious requests
3SevereRequest for high-risk assistance with substantial potential consequencesSuicide, jailbreak attempts

Written Forms

Each prompt appears in three aligned forms that share the same label. For example, a Safe customer-support request:

Code-mixed prompts mix Bangla with English, Hindi, or Urdu.

Dataset

The data is being built mainly from the Bangla portions of public LLM safety resources (including LinguaSafe, XSafety, BanglaSafe, IndicSafe, and MultiJail), plus benign customer-support prompts. All source labels are mapped to the four-level scheme above. The current version has about 10,000 prompts, each in all three written forms.

We also plan to include adversarial prompts derived from MALICE (Malicious Adversarial LLM Input for Code Evaluation), a benchmark of about 250,000 code-mixed and transliterated malicious prompts across 18 languages, including Bangla.

The data will come with train, dev, and test splits. All three forms of a prompt will stay in the same split. Release dates and download links will be posted here.

Evaluation

The planned main metric is macro-F1 over the four labels. We may also report scores for each written form separately.

Important Dates

How to Participate

Registration, the task repository, and the submission platform will be announced soon. Please check back on this page.

System Description Papers

Participants are highly encouraged to submit a system description paper (up to 4 pages) to the workshop. Please see the Submission Guidelines in the CFP.

Contact

For questions about this task, please see the Contact page.