Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

CPU- and GPU-Based Distributed Sampling in Dirichlet Process Mixtures for Large-Scale Analysis

Дата публикации: 31-05-2026 00:00:00

In the realm of unsupervised learning, Bayesian nonparametric mixture models, exemplified by the Dirichlet process mixture model (DPMM), provide a principled approach for adapting the complexity of the model to the data. Such models are particularly useful in clustering tasks where the number of clusters is unknown. Despite their potential and mathematical elegance, however, DPMMs have yet to become a mainstream tool widely adopted by practitioners. This is arguably due to a misconception that these models scale poorly as well as the lack of high-performance (and user-friendly) software tools that can handle large datasets efficiently. In this paper we bridge this practical gap by proposing a new, easy-to-use, statistical software package for scalable DPMM inference. More concretely, we provide efficient and easily-modifiable implementations for high-performance distributed sampling-based inference in DPMMs where the user is free to choose between either a multiple-machine, multiple-core, central-processing- unit (CPU) implementation (in Julia) and a multiple-stream graphics-processing-unit (GPU) implementation (in CUDA/C++). Both the CPU and GPU implementations come with a common (and optional) Python wrapper, providing the user with a single point of entry with the same interface. On the algorithmic side, our implementations leverage a leading DPMM sampler from Chang and Fisher III (2013). While Chang and Fisher III's implementation (in MATLAB/C++) used only CPU and was designed for a single multi-core machine, the packages we proposed here distribute the computations efficiently across either multiple multi-core machines or across multiple GPU streams. This leads to speedups, alleviates memory and storage limitations, and lets us fit DPMMs to significantly larger datasets and of higher dimensionality than was possible previously by either Chang and Fisher III (2013) or other DPMM methods.

Основное содержимое страницы с новостью.

Or Dinari, Raz Zamir, John W. Fisher III, Oren Freifeld

Main Article Content
Abstract

In the realm of unsupervised learning, Bayesian nonparametric mixture models, exemplified by the Dirichlet process mixture model (DPMM), provide a principled approach for adapting the complexity of the model to the data. Such models are particularly useful in clustering tasks where the number of clusters is unknown. Despite their potential and mathematical elegance, however, DPMMs have yet to become a mainstream tool widely adopted by practitioners. This is arguably due to a misconception that these models scale poorly as well as the lack of high-performance (and user-friendly) software tools that can handle large datasets efficiently. In this paper we bridge this practical gap by proposing a new, easy-to-use, statistical software package for scalable DPMM inference. More concretely, we provide efficient and easily-modifiable implementations for high-performance distributed sampling-based inference in DPMMs where the user is free to choose between either a multiple-machine, multiple-core, central-processing- unit (CPU) implementation (in Julia) and a multiple-stream graphics-processing-unit (GPU) implementation (in CUDA/C++). Both the CPU and GPU implementations come with a common (and optional) Python wrapper, providing the user with a single point of entry with the same interface. On the algorithmic side, our implementations leverage a leading DPMM sampler from Chang and Fisher III (2013). While Chang and Fisher III's implementation (in MATLAB/C++) used only CPU and was designed for a single multi-core machine, the packages we proposed here distribute the computations efficiently across either multiple multi-core machines or across multiple GPU streams. This leads to speedups, alleviates memory and storage limitations, and lets us fit DPMMs to significantly larger datasets and of higher dimensionality than was possible previously by either Chang and Fisher III (2013) or other DPMM methods.

Article Details Article Sidebar

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1BayesMultiMode: Bayesian Mode Inference in R05.4505-06-2026
2fastcpd: Fast Change Point Detection in R07.1325-07-2026
3BCDAG: An R Package for Bayesian Structure and Causal Learning of Gaussian DAGs03.8424-07-2026
4collapse: Advanced and Fast Statistical Computing and Data Transformation in R08.0631-05-2026
5Simulating Complex Cross-Sectional and Longitudinal Data Using the simDAG R Package0731-05-2026
6Dimensional Reduction for Sampled Priors and Application to Photometric Redshift Distributions05.304-08-2026
7profiling.sampling: Statistical profiler01003-01-2026
8cv: An R Package for Cross-Validating Regression Models07.315-06-2026
9balance: Deal With Biased Data Samples01027-07-2026
10WeDLM - diffusion language model01024-01-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 8.44. Источник: www.jstatsoft.org.