Blog | Documentation | Quick Start | Cookbook | SGLang | Join Slack
⭐ Star SGLang-Omni to help more builders discover open infrastructure for multimodal and speech serving!
Newsuv pip install "sglang-omni==0.1.1". [Installation] [Release notes]SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models. Its design target is multi-stage decoding: generation split across heterogeneous stages with different compute patterns, dependency structures, and resource needs. SGLang-Omni owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface, while composing with SGLang for high-performance autoregressive scheduling and model execution where applicable.
/v1/audio/speech, batch, streaming, uploaded voices./v1/audio/transcriptions. MOSS-TD supports speaker labels and timestamps (response_format=verbose_json).Additional model guides, including experimental and research-oriented paths, are available in the Cookbook.
Quick StartSGLang-Omni welcomes contributors working on inference systems, kernels, scheduling, inter-stage communication, model runners and cache efficiency, model integration, benchmarking, production deployment. Join the SGLang Slack or read the developer reference.
Organizations interested in supporting SGLang-Omni, TTS, or omni model serving can contact Chenyang Zhao at zhaochenyang@lmsys.org.
AcknowledgmentsSGLang-Omni builds on the SGLang ecosystem and on open model work from the TTS, speech, and omni-model communities. We thank the model teams, systems contributors, and partner organizations helping make open multimodal serving faster, more reliable, and easier to extend.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | omlx - LLM inference server | 0 | 10 | 06-04-2026 |
| 2 | TileRT - Tile-Based Runtime for Ultra-Low-Latency LLM Inference | 0 | 35 | 28-06-2026 |
| 3 | Memori - SQL Native Memory Layer for LLM | 0 | 10 | 27-12-2025 |
| 4 | loki-mode - Multi-agent provider agnostic framework | 0 | 43.33 | 15-02-2026 |
| 5 | semantica - Semantic Layer & Knowledge Engineering Framework | 0 | 10 | 08-02-2026 |
| 6 | inline-snapshot - Building a Robust Classifier with Stacked Generalization | 0 | 21.11 | 15-02-2026 |
| 7 | Vector Search Using Ollama for Retrieval-Augmented Generation (RAG) | 0 | 22.5 | 25-02-2026 |
| 8 | whichllm - поиск лучшей LLM модели под оборудование | 0 | 10 | 08-06-2026 |
| 9 | LLMRouter - Library for LLM Routing | 0 | 10 | 08-02-2026 |
| 10 | Semantic Caching for LLMs: FastAPI, Redis, and Embeddings | 0 | 10 | 29-04-2026 |