Research preview
Calibrated decisions, in Indian languages.
Moonim-S1 reads the text your business runs on, such as tickets, chats, call transcripts, posts, applications and an agent's notes, in 16 Indian languages and styles. It makes the call your workflow needs: where it goes, what's at risk, what happens next, and how sure it is.
See the benchmarksE-commerce support · Support ticket · Hinglish
“The refund for order 4471 has been pending for 10 days. I've emailed three times and no one has replied.”
Which team should handle this ticket?
How urgently should the team act?
Your rule→ Refunds queue, top priority
Real answers from Moonim-S1 on inputs written for our film. The rules are examples of what a team sets.
- accuracy, core track of Moonim Indic Decision Bench v0 (internal)
- 88.4
- points over billion-parameter general LLMs on the same decisions, zero-shot
- +18–25
- per message, every question answered together
- 56 ms
- calibration error in every sector and language, after a one-time calibration on your data
- ≤ 0.05
Moonim Indic Decision Bench v0 is internal. Calibration error is measured after a one-time calibration on your data.
What it is
A System 1 model: it decides, it does not write.
Moonim-S1 reads a message and a set of typed questions, and returns a calibrated probability for every option of every question. All the questions are answered in one forward pass and nothing is generated, so the answer is always one of the options you allowed. It has 4B parameters.
01
Any text
A chat, a transcript, a ticket, a form or an agent's notes.
02
Your typed questions
choice, yes / no or score, each with its own options. New schemas work zero-shot.
03
A probability per option
Every question in one forward pass. No tokens generated, nothing to parse.
One model, many decisions.
Name the questions your team needs answered and the options for each. Moonim-S1 is tuned on banking, insurance, telecom and e-commerce, and takes on new questions and new kinds of work without retraining. Each example below is from our film, in a different Indian language.
Routing and triage
Send each chat, email or ticket to the right queue, at the right priority, in any of its languages.
Hinglish · Support ticket
→ Refunds queue, top priority
Risk and compliance signals
Flag possible fraud, a likely regulator complaint or a policy breach before a person reads the case.
தமிழ் · WhatsApp chat
→ Fraud desk, card frozen
Voice and call centres
Decide on speech-to-text transcripts while the call is still going: what the caller wants, and whether to hand over.
తెలుగు · Call transcript
→ Plan desk reverts the plan, no agent needed
Agents and workflows
A fast, calibrated gate inside an automated flow: approve, hold, ask for a document, or bring in a human.
मराठी · KYC agent · step 3 of 5
→ Agent asks for the PAN card, case stays open
Moderation and quality
Check reviews, posts and replies for abuse, spam or off-policy content, written the way people actually write.
ಕನ್ನಡ · Forum post
→ Hidden, sent to the spam queue
Intake and qualification
Read applications, leads and requests in the applicant's own language and sort them before anyone opens them.
ગુજરાતી · Loan enquiry
→ MSME desk, callback today
What Moonim-S1 does
Every question, answered at once.
Ask what your workflow needs to know: which team, how urgent, is it fraud. Moonim-S1 answers all of it together, and every answer is one of the options you set, so your systems can act on it straight away.
A Hindi message, three answers
The message
मेरे account से 2,850 रुपये कट गए लेकिन UPI transaction failed दिखा रहा है, तीन दिन हो गए refund नहीं आया।
What your system gets
What is the customer's main request?
Failed transaction refund
How urgently does the bank need to act?
Soon: money or access is affected
Does it suggest fraud or a scam?
No
The same quality in every language.
Hindi, Hinglish and Tamil land within a few points of English in all four sectors we test, and beat it in three of them. Nobody has to switch to English for the model's sake.
Accuracy by sector and language
| English | Hindi | Hinglish | Tamil | |
|---|---|---|---|---|
| Banking | 91.2 | 85.6 | 86.9 | 87.2 |
| Insurance | 89.6 | 92.5 | 90.7 | 88.8 |
| Telecom | 86.4 | 86.9 | 87.5 | 90.7 |
| E-commerce | 86.9 | 88.0 | 86.9 | 88.3 |
Fast enough for every message.
All five questions about a message take 56 ms of model time (p50), answered together in one pass. A general LLM answers one question per call, so the same message takes from a third of a second to about a second.
One message, five questions, p50
Moonim-S1 4B
56 msp95 57
Indian LLM, 100B+
360 msp95 450
General LLM, 500B+
1,031 msp95 1,937
Moonim-S1: model time per message, all questions in one pass, network not included. General LLMs: one API call per question, five in parallel, timed from India with the network included.
Probabilities you can route on.
After a one-time calibration on your data, when Moonim-S1 says 90% sure it is right about 90% of the time. Set a rule such as "send to a person when it is more than 80% sure one is needed" and the rule means what it says. Asked the same questions, a 500B+ general LLM put more than 99% on a single answer in 99% of its replies.
Escalate, from the Hinglish fraud message
“Kisi ne bank se bolkar OTP maanga, maine de diya.”
routed to a person
A Hindi message, three answers
The message
मेरे account से 2,850 रुपये कट गए लेकिन UPI transaction failed दिखा रहा है, तीन दिन हो गए refund नहीं आया।
What your system gets
What is the customer's main request?
Failed transaction refund
How urgently does the bank need to act?
Soon: money or access is affected
Does it suggest fraud or a scam?
No
Next to general LLMs
Held-out messages written natively in each language, five questions each, mean accuracy: the banking and insurance slices of the core track of Moonim Indic Decision Bench v0 (internal). The general LLMs answer zero-shot through their APIs.
- Moonim-S1 4B
- General LLM, 500B+
- Indian LLM, 100B+
- Always the most common answer
English
Hindi
Hinglish
Tamil
Against other decision models
Moonim Indic Decision Bench v0 (internal). Every open model is timed on the same hardware with its own inference code; JEV is a hosted API, timed from India with the network included.
| Model | Corefour sectors, English, Hindi, Hinglish, Tamil | Independenta second, independently built test set | Extendedtwelve more languages and styles | Rewordedevery question rewritten, removing our wording advantage | p50 ms |
|---|---|---|---|---|---|
| Moonim-S1Hashtee Lab · 4B | 88.4 | 85.9 | 87.3 | 84.2 | 56 |
| JEVTypeSafe, hosted API · closed | 80.9 | 81.1 | 79.8 | 76.6 | 534 |
| Clef-flashCloudflare · 9B | 74.2 | 72.8 | 74.4 | 72.6 | 95 |
| Decider-4BMapika · 4B | 72.7 | 74.6 | 68.6 | 71.9 | 262 |
| Jet-4B v6.2quaedra · 4B | 70.8 | 69.5 | 68.0 | 70.0 | 310 |
| Decision 2.0 Nox-4Bvllm-sr · 4B | 69.4 | 68.2 | 68.6 | 69.0 | 120 |
Moonim-S1 leads every model on all four tracks.
Methodology: 4,951 native customer messages and 24,755 decisions in 16 languages and styles, in 4 tracks: core (four sectors in English, Hindi, Hinglish and Tamil), reworded questions, an independently built set, and 12 extra languages. Mean accuracy over all decisions; p50 latency per message, all questions together.
Sixteen Indian languages and styles.
Core track
- English
- हिंदी Hindi
- Hinglish
- தமிழ் Tamil
Extended track
- मराठी Marathi88.2
- বাংলা Bengali87.9
- ಕನ್ನಡ Kannada87.2
- മലയാളം Malayalam87.2
- తెలుగు Telugu87.0
- ગુજરાતી Gujarati85.0
- Banglish, Tanglish, Tenglish, Kanglish, Manglish, Marathi romanised86.1–89.1
On 1,920 messages in twelve more languages and romanised styles, Moonim-S1 scores 87.3 on average.
Available as an API.
One request per message
Send the text and all your questions together; every answer comes back in the same response. Access is on request during the research preview.
Open foundations
Built on Qwen3.5-4B-Base.
Run it on your own messages first.
We're opening Moonim-S1 to a small group of teams first, in banking, insurance, telecom, commerce and public services. Join the waitlist, or book a call to try it on a sample of your own data.

