Lythri

Hugging Face | ModelScope | GitHub | Technical Report (coming soon)

Lythri

Introduction

Lythri is a family of on-device language models built for emotional companionship. Instead of chasing math and coding scores, Lythri is trained to understand how people feel and to hold natural, multi-turn conversations, while staying small enough to run locally on a laptop or phone.

Lythri Overview

Model Total Params Active Params Base Model GGUF
Lythri-7B-A4B 7.46B 4.5B Gemma 4 E4B Lythri-7B-A4B-GGUF
Lythri-4B-A2B 4.63B 2.3B Gemma 4 E2B Lythri-4B-A2B-GGUF

Emotional Intelligence (Preliminary)

Zero-shot results.

Model GoEmotions (Macro F1) EmoBench (Acc) EQ-Bench (v2)
Gemma 4 E4B 8.00 31.39 43.19
Lythri-7B-A4B 31.54 46.83 49.41

Note: These are preliminary results from an internal evaluation that is not yet formal or complete, and they may differ slightly from the final numbers. For full details and authoritative results, including Lythri-4B-A2B, please refer to the technical report (coming soon).

General Benchmarks

Benchmark Lythri-4B-A2B Lythri-7B-A4B
Knowledge
MMLU55.2369.09
MMLU-Pro24.4238.17
ARC-E81.4083.42
ARC-C52.9960.58
Reasoning
PIQA79.4981.88
HellaSwag72.8078.29
WinoGrande68.4374.90
General
CommonsenseQA65.5277.07
SocialIQA49.8050.46
TruthfulQA MC246.1649.66
Science
OpenBookQA41.0043.60
GPQA Diamond28.7927.78
Math
GSM8K28.8162.02
MATH3.6221.28
Reading
BoolQ73.1585.32
Code / Instruction
HumanEval28.6645.12
IFEval26.4331.05

All benchmarks are evaluated with their official standard settings and in generative mode with chat template applied, reflecting real-world inference conditions. Think-tag outputs from model are stripped before answer extraction.

Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_path = "Lythri/Lythri-7B-A4B"  # or "Lythri/Lythri-4B-A2B"
tok = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
    model_path, dtype=torch.bfloat16, device_map="auto"
)

messages = [{"role": "user", "content": "My friend just lost their job and seems really down. What should I say to them?"}]
chat = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) + "<think>"
inputs = tok(chat, return_tensors="pt").to(model.device)

with torch.inference_mode():
    out = model.generate(
        **inputs,
        max_new_tokens=2048,
        do_sample=False,
        eos_token_id=[1, 106],
    )

print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Recommended Generation Config

generation_config = {
    "temperature": 0.95,
    "top_p": 0.9,
    "top_k": 64,
    "max_new_tokens": 2048,
    "repetition_penalty": 1.05,
    "do_sample": True,
    "eos_token_id": [1, 106],
}

out = model.generate(**inputs, **generation_config)

Compute

The full development of Lythri, including training and evaluation, used about 2,842 GPU hours on NVIDIA RTX 6000D GPUs.

Limitations

  • Lythri is optimized for conversation and emotional understanding, not for math, coding or complex reasoning.
  • Lythri is not a substitute for professional mental health support. If you or someone you know is in crisis, please contact local emergency services or a crisis helpline.
  • Like all language models, it can produce inaccurate or inappropriate content.

License

Lythri is built on Gemma 4 and is released under the Apache License 2.0.

Downloads last month
422
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lythri/Lythri-4B-A2B

Finetuned
(141)
this model
Quantizations
4 models

Collection including Lythri/Lythri-4B-A2B