BharatGen

Shrutam-2: LLM-Powered Multilingual Indic Speech Recognition

Shrutam-2 is a LLM based automatic speech recognition system for 12 major Indian languages. It bridges a Conformer speech encoder with a pretrained LLM decoder through a Mixture-of-Experts (MoE) projection layer, enabling high-quality, prompt-controllable transcription across diverse Indic languages.

Architecture Overview

Unlike conventional CTC/Attention ASR systems that map audio directly to text tokens, Shrutam-2 reframes speech recognition as a conditional language generation task. A speech encoder produces frame-level audio representations, which are then projected into the LLM's embedding space and fed to a frozen LLM decoder alongside a text prompt.

Languages Supported

# Language Script ISO 639-1
1 Hindi Devanagari hi
2 Marathi Devanagari mr
3 Tamil Tamil ta
4 Telugu Telugu te
5 Malayalam Malayalam ml
6 Kannada Kannada kn
7 Odia Odia or
8 Bengali Bengali bn
9 Urdu Nastaliq ur
10 Assamese Bengali as
11 Gujarati Gujarati gu
12 Punjabi Gurmukhi pa

Usage

1. Create virtual env

conda create -n shrutam2 python=3.10.14
conda activate shrutam2

2. Install dependencies

pip install torch==2.3.0 torchvision==0.18.0 torchaudio==2.3.0 --index-url https://download.pytorch.org/whl/cu118
pip install -r requirements.txt

3. Run inference

from transformers import AutoModel, AutoTokenizer
import torch

REPO_ID = "bharatgenai/Shrutam-2"

model = AutoModel.from_pretrained(REPO_ID, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
model.to("cuda" if torch.cuda.is_available() else "cpu")
model.eval()

prompt = "Transcribe speech to Hindi text."

# Single file (non-16 kHz audio is resampled automatically)
print(model.transcribe("audio.wav", prompts=prompt, tokenizer=tokenizer))

# Batch inference — one prompt per file
wavs = ["clip_hi.wav", "clip_mr.wav", "clip_ta.wav"]
prompts = [
    "Transcribe speech to Hindi text.",
    "Transcribe speech to Marathi text.",
    "Transcribe speech to Tamil text.",
]
print(model.transcribe(wavs, prompts=prompts, batch_size=2, tokenizer=tokenizer))

License

This model is released under the BharatGen non-commercial license. Please refer to the LICENSE file for detailed terms and conditions.

Shrutam 2 is developed based on the research outlined in the paper-https://arxiv.org/abs/2601.19451

Downloads last month
381
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for bharatgenai/Shrutam-2

Quantizations
1 model

Paper for bharatgenai/Shrutam-2