Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

sglang-omni - High-Performance Multi-Stage Pipeline Framework for Omni Models

Дата публикации: 27-06-2026 22:58:37



Основное содержимое страницы с новостью.


Blog | Documentation | Quick Start | Cookbook | SGLang | Join Slack

Star SGLang-Omni to help more builders discover open infrastructure for multimodal and speech serving!

News
  • [2026/08] 🚀 SGLang-Omni v0.1.1 is on PyPI. Install with uv pip install "sglang-omni==0.1.1". [Installation] [Release notes]
  • [2026/08] 🚀 TTS architecture refactor: shared pipeline state, engine construction, reference encoding, capability metadata, and vocoder scheduling. [Roadmap] [Blog]
  • [2026/06] 🔥 MOSS-TTS Local Transformer v1.5 on SGLang-Omni with native-streaming 48 kHz speech. [Blog] [Cookbook]
  • [2026/06] 🔥 Higgs Audio v3 TTS for real-time, controllable speech. [Blog] [Cookbook]

About

SGLang-Omni is a multi-stage serving runtime for omni, speech, and TTS models. Its design target is multi-stage decoding: generation split across heterogeneous stages with different compute patterns, dependency structures, and resource needs. SGLang-Omni owns the pipeline topology, stage lifecycle, inter-stage transport, model-family integration layer, and OpenAI-compatible serving surface, while composing with SGLang for high-performance autoregressive scheduling and model execution where applicable.

  • Multi-stage runtime: SGLang-Omni models generation as coordinated stages: preprocessing, encoders, autoregressive engines, talkers, decoders, vocoders, and aggregators.
  • Stage-specialized scheduling: Each stage runs behind a scheduler matched to its workload, from SGLang-backed autoregressive scheduling to lightweight preprocessing and streaming vocoder loops.
  • Transport-aware execution: A control plane coordinates requests while the relay data plane moves tensor payloads across shared-memory, NCCL, NIXL, and Mooncake backends.
  • API surface: OpenAI-compatible endpoints expose multimodal chat, speech generation, batch speech, streaming speech, uploaded voices, and transcription.

What SGLang-Omni Serves

Additional model guides, including experimental and research-oriented paths, are available in the Cookbook.

Quick Start

Community & Support

SGLang-Omni welcomes contributors working on inference systems, kernels, scheduling, inter-stage communication, model runners and cache efficiency, model integration, benchmarking, production deployment. Join the SGLang Slack or read the developer reference.

Organizations interested in supporting SGLang-Omni, TTS, or omni model serving can contact Chenyang Zhao at zhaochenyang@lmsys.org.

Acknowledgments

SGLang-Omni builds on the SGLang ecosystem and on open model work from the TTS, speech, and omni-model communities. We thank the model teams, systems contributors, and partner organizations helping make open multimodal serving faster, more reliable, and easier to extend.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1omlx - LLM inference server01006-04-2026
2TileRT - Tile-Based Runtime for Ultra-Low-Latency LLM Inference03528-06-2026
3Memori - SQL Native Memory Layer for LLM01027-12-2025
4loki-mode - Multi-agent provider agnostic framework043.3315-02-2026
5semantica - Semantic Layer & Knowledge Engineering Framework01008-02-2026
6inline-snapshot - Building a Robust Classifier with Stacked Generalization021.1115-02-2026
7Vector Search Using Ollama for Retrieval-Augmented Generation (RAG)022.525-02-2026
8whichllm - поиск лучшей LLM модели под оборудование01008-06-2026
9LLMRouter - Library for LLM Routing01008-02-2026
10Semantic Caching for LLMs: FastAPI, Redis, and Embeddings01029-04-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 43.33. Источник: pythondigest.ru.