Moonim-S1

Research preview

Calibrated decisions, in Indian languages.

Moonim-S1 reads the text your business runs on, such as tickets, chats, call transcripts, posts, applications and an agent's notes, in 16 Indian languages and styles. It makes the call your workflow needs: where it goes, what's at risk, what happens next, and how sure it is.

One email when access opens. No newsletter.

Book a call
See the benchmarks

E-commerce support · Support ticket · Hinglish

“The refund for order 4471 has been pending for 10 days. I've emailed three times and no one has replied.”

Which team should handle this ticket?

 

How urgently should the team act?

 

Your rule→ Refunds queue, top priority

Real answers from Moonim-S1 on inputs written for our film. The rules are examples of what a team sets.

accuracy, core track of Moonim Indic Decision Bench v0 (internal)
88.4
points over billion-parameter general LLMs on the same decisions, zero-shot
+18–25
per message, every question answered together
56 ms
calibration error in every sector and language, after a one-time calibration on your data
≤ 0.05

Moonim Indic Decision Bench v0 is internal. Calibration error is measured after a one-time calibration on your data.

What it is

A System 1 model: it decides, it does not write.

Moonim-S1 reads a message and a set of typed questions, and returns a calibrated probability for every option of every question. All the questions are answered in one forward pass and nothing is generated, so the answer is always one of the options you allowed. It has 4B parameters.

  1. 01

    Any text

    A chat, a transcript, a ticket, a form or an agent's notes.

  2. 02

    Your typed questions

    choice, yes / no or score, each with its own options. New schemas work zero-shot.

  3. 03

    A probability per option

    Every question in one forward pass. No tokens generated, nothing to parse.

One model, many decisions.

Name the questions your team needs answered and the options for each. Moonim-S1 is tuned on banking, insurance, telecom and e-commerce, and takes on new questions and new kinds of work without retraining. Each example below is from our film, in a different Indian language.

Routing and triage

Send each chat, email or ticket to the right queue, at the right priority, in any of its languages.

Hinglish · Support ticket

→ Refunds queue, top priority

Risk and compliance signals

Flag possible fraud, a likely regulator complaint or a policy breach before a person reads the case.

தமிழ் · WhatsApp chat

→ Fraud desk, card frozen

Voice and call centres

Decide on speech-to-text transcripts while the call is still going: what the caller wants, and whether to hand over.

తెలుగు · Call transcript

→ Plan desk reverts the plan, no agent needed

Agents and workflows

A fast, calibrated gate inside an automated flow: approve, hold, ask for a document, or bring in a human.

मराठी · KYC agent · step 3 of 5

→ Agent asks for the PAN card, case stays open

Moderation and quality

Check reviews, posts and replies for abuse, spam or off-policy content, written the way people actually write.

ಕನ್ನಡ · Forum post

→ Hidden, sent to the spam queue

Intake and qualification

Read applications, leads and requests in the applicant's own language and sort them before anyone opens them.

ગુજરાતી · Loan enquiry

→ MSME desk, callback today

One model, every decision, in 16 Indian languages: the examples above, decided by Moonim-S1 (84 s).

What Moonim-S1 does

Every question, answered at once.

Ask what your workflow needs to know: which team, how urgent, is it fraud. Moonim-S1 answers all of it together, and every answer is one of the options you set, so your systems can act on it straight away.

A Hindi message, three answers

The message

मेरे account से 2,850 रुपये कट गए लेकिन UPI transaction failed दिखा रहा है, तीन दिन हो गए refund नहीं आया।

What your system gets

What is the customer's main request?

Failed transaction refund

>99%

How urgently does the bank need to act?

Soon: money or access is affected

99%

Does it suggest fraud or a scam?

No

>99%

The same quality in every language.

Hindi, Hinglish and Tamil land within a few points of English in all four sectors we test, and beat it in three of them. Nobody has to switch to English for the model's sake.

Accuracy by sector and language

One model holds all four sectors. Mean accuracy on held-out native messages, core track of Moonim Indic Decision Bench v0 (internal).
EnglishHindiHinglishTamil
Banking91.285.686.987.2
Insurance89.692.590.788.8
Telecom86.486.987.590.7
E-commerce86.988.086.988.3

Fast enough for every message.

All five questions about a message take 56 ms of model time (p50), answered together in one pass. A general LLM answers one question per call, so the same message takes from a third of a second to about a second.

One message, five questions, p50

Moonim-S1 4B

56 msp95 57

Indian LLM, 100B+

360 msp95 450

General LLM, 500B+

1,031 msp95 1,937

Moonim-S1: model time per message, all questions in one pass, network not included. General LLMs: one API call per question, five in parallel, timed from India with the network included.

Probabilities you can route on.

After a one-time calibration on your data, when Moonim-S1 says 90% sure it is right about 90% of the time. Set a rule such as "send to a person when it is more than 80% sure one is needed" and the rule means what it says. Asked the same questions, a 500B+ general LLM put more than 99% on a single answer in 99% of its replies.

Escalate, from the Hinglish fraud message

“Kisi ne bank se bolkar OTP maanga, maine de diya.”

Should a person take over now? Yes>99%
80% · your threshold

routed to a person

Next to general LLMs

Held-out messages written natively in each language, five questions each, mean accuracy: the banking and insurance slices of the core track of Moonim Indic Decision Bench v0 (internal). The general LLMs answer zero-shot through their APIs.

  • Moonim-S1 4B
  • General LLM, 500B+
  • Indian LLM, 100B+
  • Always the most common answer

English

Moonim-S1 4B · 91.2%
91.2
General LLM, 500B+ · 70.7%
70.7
Indian LLM, 100B+ · 65.6%
65.6

Hindi

Moonim-S1 4B · 85.6%
85.6
General LLM, 500B+ · 68.0%
68.0
Indian LLM, 100B+ · 66.4%
66.4

Hinglish

Moonim-S1 4B · 86.9%
86.9
General LLM, 500B+ · 65.3%
65.3
Indian LLM, 100B+ · 64.5%
64.5

Tamil

Moonim-S1 4B · 87.2%
87.2
General LLM, 500B+ · 67.5%
67.5
Indian LLM, 100B+ · 64.5%
64.5
0255075100%

Against other decision models

Moonim Indic Decision Bench v0 (internal). Every open model is timed on the same hardware with its own inference code; JEV is a hosted API, timed from India with the network included.

ModelCorefour sectors, English, Hindi, Hinglish, TamilIndependenta second, independently built test setExtendedtwelve more languages and stylesRewordedevery question rewritten, removing our wording advantagep50 ms
Moonim-S1Hashtee Lab · 4B88.485.987.384.256
JEVTypeSafe, hosted API · closed80.981.179.876.6534
Clef-flashCloudflare · 9B74.272.874.472.695
Decider-4BMapika · 4B72.774.668.671.9262
Jet-4B v6.2quaedra · 4B70.869.568.070.0310
Decision 2.0 Nox-4Bvllm-sr · 4B69.468.268.669.0120

Moonim-S1 leads every model on all four tracks.

Methodology: 4,951 native customer messages and 24,755 decisions in 16 languages and styles, in 4 tracks: core (four sectors in English, Hindi, Hinglish and Tamil), reworded questions, an independently built set, and 12 extra languages. Mean accuracy over all decisions; p50 latency per message, all questions together.

Sixteen Indian languages and styles.

Core track

  • English
  • हिंदी Hindi
  • Hinglish
  • தமிழ் Tamil

Extended track

  • मराठी Marathi88.2
  • বাংলা Bengali87.9
  • ಕನ್ನಡ Kannada87.2
  • മലയാളം Malayalam87.2
  • తెలుగు Telugu87.0
  • ગુજરાતી Gujarati85.0
  • Banglish, Tanglish, Tenglish, Kanglish, Manglish, Marathi romanised86.1–89.1

On 1,920 messages in twelve more languages and romanised styles, Moonim-S1 scores 87.3 on average.

Available as an API.

One request per message

Send the text and all your questions together; every answer comes back in the same response. Access is on request during the research preview.

Open foundations

Built on Qwen3.5-4B-Base.

Run it on your own messages first.

We're opening Moonim-S1 to a small group of teams first, in banking, insurance, telecom, commerce and public services. Join the waitlist, or book a call to try it on a sample of your own data.

One email when access opens. No newsletter.

Book a call