Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Covariate-dependent Kato–Jones mixtures with an explicit uniform background for wind-direction regimes

Дата публикации: 22-07-2026 00:00:00

We propose a flexible and interpretable framework for circular data, motivated by wind-direction regimes. The model is a covariate-dependent finite mixture of Kato–Jones (KJ) components, reparameterized so that each component is a convex combination of a wrapped–Cauchy core and a uniform background (a WC–uniform representation). This boundary–concentration form separates regime shape from diffuse transitional mass and, at the mixture level, induces a wrapped–Cauchy mixture with an explicit uniform background. Covariates such as hour of day and wind speed enter only through a multinomial-logit link on the mixture weights, so regime incidence varies with conditions while component locations and concentrations remain globally interpretable. On the theoretical side, we establish identifiability of the covariate-dependent WC–uniform mixture via a trigonometric-moment (Toeplitz/Carathéodory–Fejér) argument under mild ordering and compactness conditions, and we derive a likelihood-ratio test for the presence of a uniform background with a chi-bar-square limit, calibrated in practice by a parametric bootstrap. For estimation we develop a block-separable sieve EM algorithm that alternates a strictly concave multinomial-logit update for the weights with closed-form moment updates for the wrapped–Cauchy parameters, guaranteeing monotone ascent of the observed likelihood and convergence to stationary points. Simulation studies with and without covariates demonstrate accurate recovery of skewed, overlapping, and switching regimes, and highlight the robustness of the WC–uniform representation to diffuse noise. Applications to NOAA buoy wind directions show that the proposed framework provides interpretable regime and uniform-background diagnostics in multi-regime, transition-rich coastal settings, while also identifying cases in which simpler von Mises mixtures or kernel-based methods provide stronger predictive likelihood.

Основное содержимое страницы с новостью.

1 Introduction

Circular measurements arise whenever the sample space is periodic: wind and ocean–current directions in geosciences; animal bearings in ecology; time–of–day events in public health and criminology; phase angles in engineering and signal processing; and joint orientations in biomechanics. Their geometry on \(\mathbb {S}^1\) requires models that are invariant to the choice of origin and that identify 0 with \(2\pi \) (Mardia and Jupp 2000; Jammalamadaka and Sengupta 2001; Fisher 1993). In many such applications, directions concentrate into a small number of regimes—for example, alternating sea/land breeze cells, preferred flight bearings, or directional traffic flows—so finite mixtures provide a natural inferential framework.

A standard baseline is the finite mixture of von Mises (vM) kernels. These mixtures are computationally convenient because each component has only a location and a concentration parameter, the density evaluation is stable through the modified Bessel normalizing constant, and standard EM implementations use simple responsibility updates together with low-dimensional component-wise updates for the circular mean and concentration. Thus, for symmetric and well-separated directional regimes, vM mixtures are typically faster, easier to initialize, and less sensitive to boundary or ordering constraints than KJ-based mixtures. A vM mixture augmented with an explicit uniform component is therefore a natural baseline, and we include both covariate-dependent and unconditional vM–uniform competitors in the empirical comparison. The limitation is that the uniform component accounts for diffuse background mass but does not change the symmetric shape of the vM kernels themselves. Several recurring features of real data strain such symmetric, light–tailed components: (i) regimes with clearly skewed or shoulder-heavy modes, especially during transitions between dominant flow patterns; (ii) diffuse background mass arising from weak forcing, unresolved physical processes, or measurement artifacts; and (iii) covariate effects (e.g., time of day, wind speed) that primarily modulate the incidence of regimes rather than their shapes (Pewsey et al. 2013; Fisher and Lee 1992). In such settings, vM–based mixtures may require additional symmetric components to approximate a single skewed or heavy-shouldered regime, thereby reducing interpretability. Asymmetric circular families offer a more direct description of skewed regimes. The Kato–Jones (KJ) class provides unimodal densities on \(\mathbb {S}^1\) with tunable skewness and peakedness, together with tractable trigonometric moments (Kato and Jones 2010, 2013). Let \(\Theta \in [-\pi ,\pi )\) and consider a KJ density with location \(\mu \in [-\pi ,\pi )\), phase \(\lambda \in [-\pi ,\pi )\), concentration \(\rho \in (0,1)\), and scale \(\gamma \in (0,\overline{\gamma }(\lambda ,\rho )]\):

$$\begin{aligned} g_{\textrm{KJ}}(\theta ;\mu ,\gamma ,\lambda ,\rho ) =\frac{1}{2\pi }\!\left\{ 1+ 2\gamma \, \frac{\cos (\theta -\mu )-\rho \cos \lambda }{1+\rho ^2-2\rho \cos (\theta -\mu -\lambda )} \right\} ,\qquad -\pi \le \theta <\pi , \end{aligned}$$

(1)

where the upper bound

$$\begin{aligned} \overline{\gamma }(\lambda ,\rho )=\frac{1-\rho ^2}{2\{1-\rho \cos \lambda \}} \end{aligned}$$

(2)

ensures nonnegativity (Kato and Jones 2013). Recent review papers (Ley and Verdebout 2017; Pewsey and García-Portugués 2021) and the traffic-flow analysis of Nagasaki et al. (2025) highlight the KJ family as a flexible alternative to classical circular models: it retains unimodality at the component level, separates parameters controlling location, concentration, skewness and kurtosis, and admits simple closed forms for trigonometric moments.

A key structural feature of the KJ density is revealed by working at the boundary \(\gamma =\overline{\gamma }(\lambda ,\rho )\). Define the boundary form

$$\begin{aligned} g^\star (\theta ;\mu ,\lambda ,\rho ):= g_{\textrm{KJ}}\!\big (\theta ;\mu ,\overline{\gamma }(\lambda ,\rho ),\lambda ,\rho \big ), \end{aligned}$$

(3)

and write \(\alpha =\gamma /\overline{\gamma }(\lambda ,\rho )\in (0,1]\). A direct algebraic rearrangement of (1)–(3) yields the convex split

$$\begin{aligned} g_{\textrm{KJ}}(\theta ;\mu ,\gamma ,\lambda ,\rho ) =\alpha \,g^\star (\theta ;\mu ,\lambda ,\rho )\;+\;(1-\alpha )\,\frac{1}{2\pi }. \end{aligned}$$

(4)

Thus, every KJ component can be written as a mixture of its boundary kernel \(g^\star \) and the uniform law. The boundary kernel behaves as a wrapped–Cauchy (WC)-type density with heavier shoulders than vM, while the uniform layer isolates genuinely diffuse mass. Summed across components, any finite mixture of KJ densities can therefore be reorganized as a finite WC mixture plus an explicit uniform background.

This boundary–uniform representation is closely related to the reparameterization proposed by Nagasaki et al. (2025) in their analysis of diurnal traffic patterns. There, the mixture is expressed in terms of KJ submodels at the boundary concentration and a single uniform component, leading to an identifiable parameterization and efficient estimation via trigonometric moments and EM. In the present work, we adopt this representation as the starting point of our modeling strategy: we treat \(g^\star \) as the “core” of each regime and the uniform distribution as a shared background that absorbs transitional or unstructured observations. This organization stabilizes estimation, prevents transitional data from inflating fitted concentrations, and keeps regime–shape parameters directly interpretable as features of physically meaningful regimes.

Mixtures of symmetric kernels need not be globally symmetric unless components are placed and weighted symmetrically. In practice, a single skewed regime (for example, a downwind coastal flow with a long return shoulder) is often represented by several symmetric components positioned asymmetrically, which complicates interpretation and inflates the apparent number of modes. This motivates the use of intrinsically asymmetric components. In our setting, KJ/WC cores avoid this proliferation: a single component can capture a skewed regime, while the explicit uniform layer absorbs diffuse contamination. More broadly, our construction belongs to a larger class of asymmetric circular models: related proposals include sine–skewed families (Abe and Pewsey 2011), wrapped–skew distributions (Pewsey 2000), and generalized von Mises variants (Gatto and Jammalamadaka 2007); see also mixture-via-wrapping approaches in Jammalamadaka and Kozubowski (2017).

In many applications, covariates primarily explain when each directional regime is active rather than how the regime looks. For wind direction, for example, the hour of day and wind speed modulate the incidence of sea/land breeze cells or synoptic regimes, while the associated mean directions and concentrations are relatively stable over time. Classical angular regression (Fisher and Lee 1992) and circular–circular link models (Kato et al. 2008) therefore motivate a mixture-of-experts specification in which predictors act on the class probabilities, and component shapes are treated as global. Building on this idea, and on the boundary representation (4), we formulate a covariate-dependent mixture in which: (i) the KJ/WC cores \(g^\star \) represent regime shapes; (ii) an explicit uniform background accounts for diffuse mass; and (iii) covariates enter only through a multinomial-logit map for the class probabilities. This parameterization reduces the confounding that arises when both component means and weights vary with predictors and aligns the model with the scientific objective of describing how external conditions affect the prevalence of distinct directional regimes.

Our empirical focus is wind direction, a canonical circular domain with extensive prior modeling via vM mixtures, nonnegative trigonometric sums (NNTS), projected–normal regressions, and wrapped distributions (Fernández-Durán 2004; Fernández-Durán and Gregorio-Domínguez 2007; Wang and Gelfand 2013). Simulations, with and without covariates, show that the WC–uniform representation recovers skewed or heavy-shouldered modes more faithfully than symmetric alternatives, while covariate-varying weights capture regime-incidence changes without spuriously shifting component centers. On public NOAA buoy records we find a consistent dichotomy: in open–ocean months with a single tight mode and negligible diffuse mass, simple vM mixtures can be marginally more efficient; in coastal months with multi-regime structure, skewness, and transitional diffusion, the WC–uniform formulation provides clearer regime and background diagnostics, although the strongest predictive likelihood may be achieved by more flexible kernel-based or von Mises competitors.

Outline. Section 2 introduces the covariate-dependent KJ mixture, records the WC–uniform map, and states identifiability. Section 3 develops a sieve EM algorithm with closed–form WC moment updates and a boundary likelihood-ratio test for the uniform background. Section 4 studies finite–sample performance with and without covariates. Section 5 analyzes NOAA buoy records. Section 6 compares against vM–based competitors. Section 7 concludes with a summary, limitations, and possible extensions.

2 Model: covariate-dependent mixtures of Kato–Jones submodels with uniform background
2.1 Set-up and notation

Let the sample space be the circle \(\mathbb {S}^1=[-\pi ,\pi )\) with Lebesgue measure \(d\theta \), and let \(\mathcal {X}\) denote the covariate space. We use the Kato–Jones density \(g_{\textrm{KJ}}\) in (1) with boundary concentration (2) and boundary form \(g^\star \) in (3). The boundary–uniform convex split (4) will be used repeatedly. Background on directional statistics and KJ properties appears in Mardia and Jupp (2000); Jammalamadaka and Sengupta (2001); Kato and Jones (2010, 2013).

2.2 Covariate-dependent mixture with explicit uniform background

Fix an integer \(m\ge 1\). The model has \(m+1\) mixture components: m structured boundary KJ components and one explicit uniform background component. We denote the uniform weight by \(\pi _u(x)\) throughout, so that \(\pi _u(x)\equiv \pi '_{m+1}(x)\). For \((\Theta ,X)=(\theta ,x)\in \mathbb {S}^1\times \mathcal {X}\), define the conditional density

$$\begin{aligned} f(\theta \mid x;\Psi )\;=\;\sum _{k=1}^{m}\pi '_k(x)\, g^{\star }\!\bigl (\theta ;\mu _k,\lambda _k,\rho _k\bigr ) \;+\;{\pi _u(x)}\,\frac{1}{2\pi }, \end{aligned}$$

(5)

where \(\Psi \) collects global component parameters

$$ (\mu _k,\lambda _k)\in [-\pi ,\pi ),\qquad \rho _k\in (0,1),\qquad k=1,\dots ,m, $$

and measurable maps \(x\mapsto \pi '_k(x)\in [0,1]\) (\(k=1,\dots ,m\)) and \(x\mapsto \pi _u(x)\in [0,1]\) satisfying \(\sum _{k=1}^{m}\pi '_k(x)+\pi _u(x)=1\) for all \(x\in \mathcal {X}\). Classes 1 : m are boundary KJ components; the remaining class is the explicit uniform background. We refer to (5) as the covariate-dependent mixture of KJ submodels with uniform background (CD–MoKJ).

A classical KJ mixture without an explicit background has the form

$$ f_{\textrm{KJ}}(\theta \mid x) =\sum _{k=1}^m \pi _k(x)\, g_{\textrm{KJ}}\!\bigl (\theta ;\mu _k,\gamma _k(x),\lambda _k,\rho _k\bigr ), $$

with weights \(\pi _k(x)\ge 0\), \(\sum _{k=1}^m\pi _k(x)=1\) and scales \(\gamma _k(x)\in \bigl (0,\overline{\gamma }(\lambda _k,\rho _k)\bigr ]\). Applying the split (4) componentwise yields a representation of the form (5) with

$$\begin{aligned} \pi '_k(x)\;=\;\pi _k(x)\,\frac{\gamma _k(x)}{\overline{\gamma }(\lambda _k,\rho _k)}, \quad k=1,\dots ,m, \quad {\pi _u(x)}\;=\;\sum _{k=1}^m \pi _k(x)\Bigl \{1-\frac{\gamma _k(x)}{\overline{\gamma }(\lambda _k,\rho _k)}\Bigr \}. \end{aligned}$$

(6)

Thus the CD–MoKJ parameterization is at least as general as the standard KJ mixture, while making the uniform background explicit. In line with mixture-of-experts designs for angular data, covariates act only on the class probabilities \(\{\pi '_1(x),\ldots ,\pi '_m(x),\pi _u(x)\}\); the component shapes \((\mu _k,\lambda _k,\rho _k)\) do not depend on x (Fisher and Lee 1992; Jacobs et al. 1991; Yuksel et al. 2012).

2.3 Standing assumptions and measurability

Assume compactness and positivity: there exist \(0<\underline{r}<\overline{r}<1\) and \(0<\underline{w}<1\) such that

$$ \rho _k\in [\underline{r},\overline{r}]\quad (k=1,\dots ,m),\qquad \pi '_k(x)\in [\underline{w},1)\quad (k=1,\dots ,m),\qquad {\pi _u(x)\in [0,1),} $$

and \(\sum _{k=1}^{m}\pi '_k(x)+\pi _u(x)=1\) for all \(x\in \mathcal {X}\). Labels are fixed by the ordering \(0<\rho _m<\cdots<\rho _1<1\). The covariate law \(P_X\) has support with nonempty interior. Each map \(x\mapsto \pi '_k(x)\) is Borel measurable. Continuity of \(g^\star \) and linearity of (5) imply that \((\theta ,x,\Psi )\mapsto f(\theta \mid x;\Psi )\) is Borel measurable and dominated by Lebesgue measure on \(\mathbb {S}^1\).

2.4 Identifiability with covariates

For \(p\in \mathbb {Z}\) define the trigonometric moments

$$ M_p(x):=\int _{-\pi }^{\pi } e^{\textrm{i}p\theta }\,f(\theta \mid x;\Psi )\,d\theta . $$

For \(p\ne 0\) the uniform term in (5) vanishes, and the boundary KJ submodel has analytic moments

$$ \int _{-\pi }^{\pi } e^{\textrm{i}p\theta }\,g^\star (\theta ;\mu ,\lambda ,\rho )\,d\theta \;=\;c_p(\lambda ,\rho )\,e^{\textrm{i}p\mu },\qquad p\in \mathbb {N}, $$

with \(c_p(\lambda ,\rho )\ne 0\) real-analytic on \((-\pi ,\pi )\times (0,1)\) (Kato and Jones 2013). Hence, for fixed x,

$$ M_p(x)=\sum _{k=1}^m \pi '_k(x)\,c_p\!\bigl (\lambda _k,\rho _k\bigr )\,e^{\textrm{i}p\mu _k}\qquad (p\in \mathbb {N}), $$

so \(\{M_p(x)\}_{p\ge 0}\) is a finite sum of complex exponentials in p.

Theorem 2.1

(Identifiability of CD–MoKJ) Fix \(m\ge 1\). Suppose two models of the form (5) satisfy the conditions in Sect.2.3 and the same ordering of \(\{\rho _k\}\). If

$$\begin{aligned} f(\theta {\mid }x;\Psi ) = \tilde{f}(\theta {\mid }x;\tilde{\Psi })\quad for\,P_{X} - almost\,every\,x\,and\,all\,\theta \in {\mathbb {S}}^{1} , \end{aligned}$$

then, up to a common permutation of component labels, the global parameters \((\mu _k,\lambda _k,\rho _k)\) and the weight functions \(\{\pi '_k(\cdot )\}_{k=1}^{m}\) and \(\pi _u(\cdot )\) coincide; in particular, \(\pi _u(\cdot )\) is identified.

Proof

Fix x in a full \(P_X\)-measure set with \(f(\cdot \mid x;\Psi )=\tilde{f}(\cdot \mid x;\widetilde{\Psi })\). Define

$$ \alpha _k(x):=\pi '_k(x)\,c_1\!\bigl (\lambda _k,\rho _k\bigr )>0,\qquad z_k:=e^{\textrm{i}\mu _k}. $$

Then \(M_p(x)=\sum _{k=1}^m \alpha _k(x)\,z_k^{\,p}\) for \(p\ge 1\). Let \(T_m(x)\) be the \((m{+}1)\times (m{+}1)\) Hermitian Toeplitz matrix with entries \([T_m(x)]_{ij}=M_{j-i}(x)\) (\(0\le i,j\le m\), \(M_{-q}=\overline{M_q}\)). Because \(f(\cdot \mid x)\) is a density, \(T_m(x)\succeq 0\). By the Carathéodory–Fejér theorem (Grenander and Szegö 1958, Ch. 3), if \(\operatorname {rank}T_m(x)=m\)—which holds here since the ordering and compactness prevent coalescence, \(c_1(\lambda ,\rho )>0\) on \((-\pi ,\pi )\times (0,1)\) (Kato and Jones 2013), and \(\pi '_k(x)>0\)—then \(\{z_k\}\) and \(\{\alpha _k(x)\}\) are uniquely determined up to permutation.

For \(p=2,\dots ,2m\),

$$ \frac{M_p(x)}{M_1(x)}=\sum _{k=1}^m \beta _k(x)\, \frac{c_p\!\bigl (\lambda _k,\rho _k\bigr )}{c_1\!\bigl (\lambda _k,\rho _k\bigr )}\, z_k^{\,p-1},\qquad \beta _k(x):=\frac{\alpha _k(x)}{\sum _{\ell =1}^m\alpha _\ell (x)}\in (0,1). $$

Real-analyticity and positivity of \(c_p\) imply local injectivity of \((\lambda ,\rho )\mapsto (c_p/c_1)_{p=2}^{2m}\) on a dense open set, so \((\lambda _k,\rho _k)\) are identified componentwise once \(\{z_k\}\) are known. Finally,

$$ \pi '_k(x)=\frac{\alpha _k(x)}{c_1\!\bigl (\lambda _k,\rho _k\bigr )}\in (0,1),\qquad {\pi _u(x)=1-\sum _{k=1}^m \pi '_k(x),} $$

which identifies the uniform background.

The same argument applies to the tilde-parameterization, and equality of the moment sequences forces equality of the parameters up to a common permutation. The conclusion holds \(P_X\)-almost surely. \(\square \)

2.5 Interpretation in practice

The CD–MoKJ parameterization (5) is identifiable and numerically stable, and the mapping (6) shows there is no loss of generality relative to classical KJ mixtures. The explicit uniform background isolates diffuse transitional mass, while covariate-dependent weights reallocate regime incidence over x without spuriously shifting component centers.

The connection with Nagasaki et al. (2025) is structural but not identical in emphasis. Their work shows that KJ mixtures can be expressed through boundary KJ submodels together with a uniform component, leading to an identifiable and computationally convenient formulation for applied circular mixture modeling. We use the same boundary–uniform algebra as the starting point, but we build it into a covariate-dependent mixture-of-experts model in which the class probabilities vary with predictors while the regime shapes remain globally interpretable. Thus, in the present formulation, the uniform component is not only part of a reparameterization; it is treated as an explicit background layer whose magnitude can be estimated, diagnosed, and tested.

This distinction is important for wind-direction applications. In the formulation used here, the WC/KJ cores represent structured directional regimes, the uniform layer represents diffuse or transitional observations, and the multinomial-logit weights describe how the incidence of these regimes changes with covariates such as hour of day and wind speed. The framework therefore separates three inferential questions: the shape of each directional regime, the amount of unstructured background mass, and the covariate-dependent prevalence of each component. Compared with an unconditional KJ-mixture fit, this organization makes the role of diffuse mass more transparent and provides direct diagnostics, including the boundary likelihood-ratio test for the uniform layer.

In multi-regime, skewed, or contaminated settings, the heavier-shouldered KJ core provides robustness; the empirical sections quantify these gains. For model selection we use information-theoretic criteria grounded in Kullback–Leibler risk (Kullback and Leibler 1951).

3 Estimation: Sieve–EM for CD–MoKJ
3.1 Finite-dimensional sieves and parameterization

Let \(\{\phi _r(\cdot )\}_{r=1}^{q}\) be a fixed, linearly independent real basis on \(\mathcal {X}\) (e.g., B-splines or low-degree polynomials) and write \(\phi (x)=(\phi _1(x),\ldots ,\phi _q(x))^\top \). We approximate the covariate effect on the class probabilities with a fixed-q sieve:

$$ \eta _k(x)=\beta _k^\top \phi (x),\qquad \pi '_k(x)=\frac{\exp \{\eta _k(x)\}}{\sum _{\ell =1}^{m+1}\exp \{\eta _\ell (x)\}},\quad k=1,\ldots ,m{+}1. $$

For notational consistency with Sect. 2.2, we write \(\pi _u(x)=\pi '_{m+1}(x)\) for the uniform component. Component shapes are global: for \(k=1,\ldots ,m\),

$$ \mu _k\in [-\pi ,\pi ),\qquad \lambda _k\in [-\pi ,\pi ),\qquad \rho _k=h(\xi _k),\quad h(u)=\frac{1}{1+e^{-u}},\ \xi _k\in \mathbb {R}, $$

so that \(0<\rho _k<1\). Collect the unknown parameters in

$$ \Xi =\{\beta _2,\ldots ,\beta _{m+1};\ \mu _1,\ldots ,\mu _m;\ \lambda _1,\ldots ,\lambda _m;\ \xi _1,\ldots ,\xi _m\}. $$

For identifiability we take component \(k=1\) as the softmax baseline and set \(\beta _1\equiv 0\), and we constrain all free coefficients to a large compact box. Label ordering \(0<\rho _m<\cdots<\rho _1<1\) is enforced via the cumulative-logit reparameterization in Remark 3.1. Throughout, the boundary KJ kernel \(g^\star \) is the one defined in (3); we do not redefine it here.

3.2 Observed and complete-data likelihoods

Given i.i.d. observations \(\{(\Theta _j,X_j)\}_{j=1}^n\) with angles \(\Theta _j\in [-\pi ,\pi )\), introduce latent indicators \(Z_{jk}\in \{0,1\}\) for \(k=1,\ldots ,m\) and \(Z_{j,u}\in \{0,1\}\) for the uniform class, with \(\sum _{k=1}^{m}Z_{jk}+Z_{j,u}=1\). The complete-data log-likelihood is

$$ \ell _c(\Xi )=\sum _{j=1}^n\sum _{k=1}^{m} Z_{jk}\Big \{\log \pi '_k(X_j)+\log g^\star \big (\Theta _j;\mu _k,\lambda _k,\rho _k\big )\Big \} +\sum _{j=1}^n {Z_{j,u}\log \pi _u(X_j)}. $$

The observed log-likelihood is

$$\begin{aligned} \ell (\Xi )=\sum _{j=1}^n \log \!\Bigg (\sum _{k=1}^{m}\pi '_k(X_j)\, g^\star \big (\Theta _j;\mu _k,\lambda _k,\rho _k\big )\;+\;{\pi _u(X_j)}\,\frac{1}{2\pi }\Bigg ). \end{aligned}$$

This parameterization uses only the boundary form \(g^\star \) and the explicit uniform background; recovery of the classical KJ mixture parameters is via (6) in Sect. 2.

3.3 A monotone sieve–EM algorithm

At iteration t with parameter \(\Xi ^{(t)}\), define responsibilities

$$ w_{jk}^{(t)}:=\mathbb {P}^{(t)}(Z_{jk}=1\mid \Theta _j,X_j),\qquad N_k^{(t)}=\sum _{j=1}^n w_{jk}^{(t)}. $$

For \(k\le m\),

$$ w_{jk}^{(t)}=\frac{\pi _k'^{(t)}(X_j)\,g^\star \!\big (\Theta _j;\mu _k^{(t)},\lambda _k^{(t)},\rho _k^{(t)}\big )}{\sum _{\ell =1}^{m}\pi _\ell '^{(t)}(X_j)\,g^\star \!\big (\Theta _j;\mu _\ell ^{(t)},\lambda _\ell ^{(t)},\rho _\ell ^{(t)}\big )+{\pi _u^{(t)}(X_j)}\,\tfrac{1}{2\pi }}, $$

and for the uniform class

$$ {w_{j,u}^{(t)}}=\frac{{\pi _u^{(t)}(X_j)}\,\tfrac{1}{2\pi }}{\sum _{\ell =1}^{m}\pi _\ell '^{(t)}(X_j)\,g^\star \!\big (\Theta _j;\mu _\ell ^{(t)},\lambda _\ell ^{(t)},\rho _\ell ^{(t)}\big )+{\pi _u^{(t)}(X_j)}\,\tfrac{1}{2\pi }}. $$

The EM surrogate is

$$ Q(\Xi ;\Xi ^{(t)})=\sum _{j=1}^n\sum _{k=1}^{m} w_{jk}^{(t)}\log \pi '_k(X_j) +\sum _{j=1}^n {w_{j,u}^{(t)}\log \pi _u(X_j)} +\sum _{j=1}^n\sum _{k=1}^{m} w_{jk}^{(t)}\log g^\star \!\big (\Theta _j;\mu _k,\lambda _k,\rho _k\big ). $$

Algorithm 1

Algorithm 1: Sieve–EM for CD–MoKJ

Remark 3.1

(Practical enforcement of ordering) Let \(\tilde{\rho }_k=h(\xi _k)\in (0,1)\) for \(k=1,\ldots ,m\). Set \(\rho _m=\tilde{\rho }_m\) and recursively define

$$ \rho _{k-1}=\rho _k+\{1-\rho _k\}\tilde{\rho }_{k-1},\quad k=m,\ldots ,2, $$

which yields strict ordering \(0<\rho _m<\cdots<\rho _1<1\) without post-hoc label switching.

Remark 3.2

(Initialization, stability, and complexity). A practical two-stage initialization is: (i) fit a coarse m-component von Mises mixture (or NNTS) ignoring covariates to obtain initial centers and concentrations; (ii) compute soft labels from this fit and regress them on \(\phi (x)\) to initialize \(\{\beta _k\}\), while \((\mu _k,\lambda _k,\xi _k)\) are initialized from the unconditional fit. Small ridge penalties on \(\{\beta _k\}\) help avoid complete separation in early iterations and can be removed once \(\{w_{jk}\}\) stabilize. Each EM iteration costs O(nmq) operations, and memory usage is O(nm) for the responsibilities. In practice, the main computational burden comes from multi-start fitting, BIC searches over m, and bootstrap refitting rather than from a single EM iteration. Since the number of free parameters is \(d_m=m(q+3)\), we keep both the number of components and the covariate basis dimension modest, and we inspect the effective class sizes \(N_k=\sum _{j=1}^n \widehat{w}_{jk}\) to avoid retaining components supported by only a small fraction of the sample. For the wind-direction examples considered below, this leads naturally to small grids for m and low-dimensional covariate bases.

3.4 Monotone ascent and limit points

On the compact feasible set (baseline fixed, coefficient boxes), the EM map is closed and each M-step maximizes \(Q(\cdot ;\Xi ^{(t)})\). The EM inequality gives \(\ell (\Xi ^{(t+1)})\ge \ell (\Xi ^{(t)})\); accumulation points satisfy KKT conditions for the constrained problem.

Proposition 3.1

(EM monotonicity and stationary limit points) Let \(\{\Xi ^{(t)}\}\) be produced by Algorithm 1 from any \(\Xi ^{(0)}\) in the feasible set. Then \(\ell (\Xi ^{(t+1)})\ge \ell (\Xi ^{(t)})\) for all t, with equality only at fixed points. Every accumulation point \(\Xi ^\star \) is a stationary point of \(\ell \) on the feasible set.

Proof

Standard EM arguments (Wu 1983). The multinomial block is strictly concave, and the component blocks are smooth over a compact set, ensuring closedness of the map and upper semicontinuity of Q.

\(\square \)

Parameter count for BIC. With \(\beta _1\) fixed as baseline, the softmax contributes mq free coefficients; \((\mu ,\lambda ,\rho )\) contribute 3m scalars. Hence

$$\begin{aligned} d_m=m(q+3). \end{aligned}$$

(7)

This count gives a useful practical diagnostic: increasing either the number of regimes m or the covariate basis dimension q increases the model dimension linearly, but it can also reduce the effective information available for each component. We therefore recommend selecting m over a small interpretable grid using BIC or cross-validated likelihood, keeping q modest unless there is clear predictive evidence for additional covariate terms, and checking the effective class sizes \(N_k=\sum _{j=1}^n\widehat{w}_{jk}\) after fitting. With limited samples, smaller values of m and simple covariate effects are preferable, since weakly supported components can be difficult to distinguish from the uniform background or from low-concentration WC tails.

3.5 Model selection via BIC

For \(m\in \{1,\ldots ,m_{\max }\}\), let \(\widehat{\Xi }_m\) be a (local) maximizer returned by Algorithm 1 and set \(\widehat{\ell }_m=\ell (\widehat{\Xi }_m)\). Define

$$\begin{aligned} \textrm{BIC}(m)\;=\;-2\,\widehat{\ell }_m\;+\;d_m\log n, \end{aligned}$$

(8)

with \(d_m\) in (7). In a correctly specified fixed-q sieve with true order \(m_0\) and standard mixture regularity (identifiability, interior nonzero weights, nonsingular Fisher information modulo labels), BIC consistently selects \(m_0\).

Proposition 3.2

(BIC consistency in the fixed-q sieve) Assume the data are generated by a CD–MoKJ model within the fixed-q sieve with \(m_0\) components, and \(\widehat{\Xi }_m\) are consistent local maximizers. Then \(\arg \min _{1\le m\le m_{\max }}\textrm{BIC}(m)\xrightarrow {p} m_0\) as \(n\rightarrow \infty \).

Proof

As in Keribin (2000): for \(m<m_0\), the Kullback–Leibler gap induces an O(n) loss that dominates the \(\log n\) penalty; for \(m>m_0\), identifiability bounds the likelihood gain to \(O_p(1)\) while the penalty increases by \((d_m-d_{m_0})\log n\). \(\square \)

3.6 Testing for a covariate-dependent uniform background

To assess the need for a uniform layer, consider

$$ H_0:\ {\pi _u(x)}\equiv 0 \quad \text {vs.}\quad H_1:\ {\pi _u(x)}>0\ \text {for some }x. $$

In the sieve, \(H_1\) adds a q-dimensional coefficient vector \(\beta _u\equiv \beta _{m+1}\) for the uniform class relative to the fixed baseline. Let \(\widehat{\ell }^{(m)}\) be the maximized log-likelihood under the reduced model (no uniform) and \(\widehat{\ell }^{(m+1)}\) under the full model. The likelihood-ratio statistic is

$$\begin{aligned} \Lambda _n\;=\;2\big \{\widehat{\ell }^{(m+1)}-\widehat{\ell }^{(m)}\big \}. \end{aligned}$$

(9)

Because \(H_0\) places the added weights on the boundary of the parameter space, \(\Lambda _n\) has a chi-bar-square limit with weights determined by the covariance of the score for the added logits.

Proposition 3.3

(Chi-bar-square limit at the boundary) Under \(H_0\) and standard regularity (identifiability, interiority of the other weights, nonsingular information for the active block), \(\Lambda _n \Rightarrow \sum _{j=0}^{q}\omega _j\,\chi _j^2\), where \(\{\omega _j\}\) are determined by the Gaussian limit of the score and the corresponding information block (Self and Liang 1987; Silvapulle and Sen 2005). If the block is whitened to near-spherical form, then \(\omega _j\approx \left( {\begin{array}{c}q\\ j\end{array}}\right) 2^{-q}\).

Calibration. (i) Whiten the added block by the observed information at \(\widehat{\Xi }^{(m)}\) and use \(\omega _j=\left( {\begin{array}{c}q\\ j\end{array}}\right) 2^{-q}\); or (ii) employ a parametric bootstrap under the fitted null: simulate from the m-component model without the uniform layer, refit both models, and recompute \(\Lambda _n\). In the finite-sample simulations and data analyses below, we use the bootstrap calibration when reporting evidence for a uniform component, because the boundary null can be sensitive to overlap between low-concentration WC tails and the explicit uniform layer.

Remark 3.3

(Asymptotics of sieve estimators) With fixed q and interior parameters, the sieve estimator is regular: linear functionals of \((\beta ,\mu ,\lambda ,\rho )\) are asymptotically normal at rate \(\sqrt{n}\). Pointwise inference on \(x\mapsto \pi '_k(x)\) follows by the delta method. Under \(H_0\) in Sect.3.6, boundary corrections via the chi-bar-square limit are required for functionals involving the uniform class.

4 Simulation study

We conduct three complementary experiments. First, a baseline study without covariates mirrors the empirical configuration in Sect. 5, probing finite-sample accuracy, overlap sensitivity, and the benefit of an explicit uniform background. Second, we consider null-uniform designs with \(\pi _u=0\) to evaluate the finite-sample behavior of the boundary likelihood-ratio test in Sect. 3.6. Third, a compact covariate-dependent design varies only the class probabilities with a univariate covariate X, illustrating how covariates improve selection and identifiability while component shapes remain global. The simulation designs are intentionally low-dimensional so that the effects of overlap, background mass, and covariate-varying regime incidence can be isolated; richer covariate bases enter through the same softmax block and increase the model dimension according to (7).

Throughout, the wrapped–Cauchy (WC) kernel is

$$\begin{aligned} \textrm{WC}(\theta \mid \phi ,\rho )\;=\;\frac{1-\rho ^2}{2\pi \{1+\rho ^2-2\rho \cos (\theta -\phi )\}}, \qquad \theta \in [0,2\pi ),\ \phi \in [0,2\pi ),\ \rho \in (0,1). \end{aligned}$$

(10)

While (10) is written on \([0,2\pi )\), all simulated and estimated angles are wrapped to \([-\pi ,\pi )\) for consistency with Sects. 13; the two parameterizations are equivalent under \(2\pi \)-periodic wrapping.

4.1 Mixtures without covariates

Design. We work with the WC–uniform reparameterization of the KJ mixture. For each component k,

$$ \phi _k=(\mu _k+\lambda _k)\bmod 2\pi ,\qquad \rho _k\in (0,1),\qquad \pi '_k\in (0,1), $$

and the uniform weight \(\pi _u\in (0,1)\) satisfies \(\sum _{k=1}^m \pi '_k+\pi _u=1\). The baseline (two components plus a small uniform share) mirrors Sect. 5:

$$\begin{aligned} (\mu _1,\lambda _1,\rho _1,\pi '_1)&=(2.7572,\,5.3136,\,0.7266,\,0.4536),\\ (\mu _2,\lambda _2,\rho _2,\pi '_2)&=(4.0107,\,1.1895,\,0.1970,\,0.4825),\quad \pi _u=0.0639, \end{aligned}$$

with \(\phi _k=(\mu _k+\lambda _k)\bmod 2\pi \). Simulations generate \(\theta \in [0,2\pi )\), immediately wrapped to \([-\pi ,\pi )\) before evaluation.

Estimator and implementation. Maximum likelihood is computed by EM under the WC–uniform parameterization with closed-form moment updates for the component block:

$$ R_k=\frac{\sum _{i=1}^n \tau _{ik}\,e^{\textrm{i}\theta _i}}{\sum _{i=1}^n \tau _{ik}},\,\, \widehat{\phi }_k=\textrm{Arg}(R_k)\!\!\pmod {2\pi },\,\, \widehat{\rho }_k=|R_k|,\,\, \widehat{\pi }'_k=\frac{1}{n}\sum _{i=1}^n \tau _{ik},\,\, \widehat{\pi }_u=\frac{1}{n}\sum _{i=1}^n \tau _{i,u}, $$

where \(\tau _{ik}\) and \(\tau _{i,u}\) are the E-step responsibilities. Angle errors for \(\phi _k\) use circular wrapping to \([-\pi ,\pi )\). We employ 10 random restarts, box constraints \(\rho _k\in (10^{-3},0.999)\) and \(\pi '_k,\pi _u\in (10^{-8},1)\), and terminate when the relative change in the observed log-likelihood drops below \(10^{-7}\). Estimating trigonometric moments (ETM) initializations based on low-order trigonometric moments (up to order 2m) produce similar results.

Sample sizes and metrics. We take \(n\in \{100,500,1000,3000,5000,10000\}\) with replicate counts \(\{8,8,6,4,3,3\}\), respectively. The number of replicates is reduced at the largest sample sizes because each replicate requires a full multi-start EM fit, and the computational cost increases substantially with n. The simulation summaries are intended to document broad finite-sample trends rather than to support fine Monte Carlo comparisons; the main conclusions are therefore based on the systematic decay of bias and RMSE across increasing sample sizes. Angular error is

$$ e_\phi (\widehat{\phi },\phi )=\big ((\widehat{\phi }-\phi +\pi )\bmod 2\pi \big )-\pi . $$

We report Bias and RMSE for \((\phi _1,\phi _2,\rho _1,\rho _2,w_1,w_2,\pi _u)\), where \(w_k=\pi '_k\) denotes the fitted weight of the kth WC core and \(\pi _u\) denotes the explicit uniform weight, and a generalized MSE summary based on the determinant of the empirical covariance matrix \(\widehat{\Sigma }\) of \((\phi _1,\phi _2,\rho _1,\rho _2,w_1,w_2)\):

$$ \widehat{\textrm{GMSE}}=|\det (\widehat{\Sigma })|. $$

Results and interpretation. Tables 1 and 2 show the expected RMSE decay with n, and Table 3 displays a rapid decrease of GMSE. The second component \((\phi _2,\rho _2)\) remains noisier at small n due to its low concentration (\(\rho _2\approx 0.20\)), which shortens the resultant and inflates phase dispersion. Weights \((w_1,w_2)\) and \(\pi _u\) are accurately recovered; the small-sample upward bias of \(\pi _u\) should be interpreted in light of the near-boundary nature of this parameter. In the present design the true uniform share is small, \(\pi _u=0.0639\), so finite-sample likelihood maximization is affected by both the lower boundary at \(\pi _u=0\) and the overlap between the uniform layer and low-concentration WC tails. Consequently, some weakly structured observations may be allocated to the uniform class, producing a positive bias that decreases as information increases. When diffuse background mass is genuinely present, the explicit uniform layer prevents transitional observations from spuriously sharpening component concentrations. The next experiment examines the complementary null case in which no uniform background is present and evaluates whether the boundary test avoids selecting the uniform layer without sufficient evidence.

Table 1 Bias and RMSE (MLE, WC–uniform EM).

Full size table

Table 2 Bias and RMSE (MLE, WC–uniform EM).

Full size table

Table 3 Generalized MSE for \((\phi _1,\phi _2,\rho _1,\rho _2,w_1,w_2)\).

Full size table

4.2 Null-uniform simulations and finite-sample calibration of the boundary test

We next examine the case in which the explicit uniform component is absent. This directly evaluates the finite-sample behavior of the likelihood-ratio test in Sect. 3.6 and highlights the boundary issue associated with testing \(\pi _u=0\). We consider two null-uniform designs. The first uses the same two WC cores as Sect. 4.1, but sets the true uniform weight to zero and renormalizes the two WC weights. This is a challenging null case because the second WC component has low concentration, \(\rho _2=0.1970\), and therefore produces diffuse tails. The second design keeps the same component locations and weights but increases the second concentration to \(\rho _2=0.5000\), giving a moderately separated null case:

$$ \pi _u=0,\qquad w_1=0.4846,\qquad w_2=0.5154, $$

obtained by renormalizing the original two WC-core weights after removing the uniform component. For each simulated data set, we fit the reduced two-component WC model without a uniform component and the full WC–uniform model with the same two WC cores plus an explicit uniform layer. The likelihood-ratio statistic is computed as in (9). Since the null value \(\pi _u=0\) lies on the boundary of the parameter space, the test is calibrated by a parametric bootstrap under the fitted null model: data are generated from the fitted no-uniform model, both the null and full models are refitted, and the bootstrap distribution of the likelihood-ratio statistic is used to compute the p-value.

Table 4 Null-uniform simulations for the boundary likelihood-ratio test of \(H_0:\pi _u=0\).

Full size table

Table 4 shows that the finite-sample behavior of the uniform-background test depends on the separation between the WC tails and the uniform layer. In the moderate-concentration null design, the fitted values of \(\widehat{\pi }_u\) remain small and the rejection rates are closer to the nominal level than in the more diffuse design. In the low-concentration null design, rejection rates are more variable, especially at \(n=500\) and \(n=3000\), reflecting the difficulty of distinguishing an explicit uniform layer from the diffuse tail of a low-concentration WC component. These results emphasize that the uniform component should not be interpreted from \(\widehat{\pi }_u\) alone. The bootstrap-calibrated likelihood-ratio test provides a more appropriate finite-sample diagnostic, while small positive estimates of \(\pi _u\) should be treated cautiously when one structured component is itself highly diffuse.

4.3 Covariate-dependent mixtures (weights varying with a covariate)

We next let only the weights vary with a univariate covariate \(X\in [0,1]\), keeping component shapes fixed:

$$ f(\theta \mid x)=\sum _{k=1}^2 \pi '_k(x)\,\textrm{WC}(\theta \mid \phi _k,\rho _k) \;+\;\pi '_3(x)\,\frac{1}{2\pi },\qquad \sum _{k=1}^3 \pi '_k(x)=1, $$

with the kernel given by (10). Shapes match the no-covariate study via \(\phi _k=(\mu _k+\lambda _k)\bmod 2\pi \):

$$ \phi _1=(2.7572+5.3136)\bmod 2\pi ,\ \rho _1=0.7266; \qquad \phi _2=(4.0107+1.1895)\bmod 2\pi ,\ \rho _2=0.1970. $$

Weights follow a softmax on \(\phi (x)=(1,x,(x-\tfrac{1}{2})^3)\):

$$ \big (\pi '_1(x),\pi '_2(x),\pi '_3(x)\big )=\textrm{softmax}\big (\beta _1^\top \phi (x),\ \beta _2^\top \phi (x),\ \beta _3^\top \phi (x)\big ), $$

with

$$ \beta _1=(0.3,-0.4,0.2)^\top ,\quad \beta _2=(0.1,0.4,-0.1)^\top ,\quad \beta _3=(-2.7,0.2,0)^\top . $$

Draw \(X\sim \textrm{Unif}(0,1)\) and then sample \(\Theta \mid X=x\) from the model.

Estimation adopts the weights-only CD–MoKJ sieve–EM specialized to this setting: E-step responsibilities under the current softmax; a multinomial-logit M-step for \(\{\pi '_k(x)\}\) with responsibilities as observation weights; and responsibility-weighted WC moment updates for \((\phi _k,\rho _k)\). The uniform class is the \((m{+}1)\)th softmax category. We compare \(m=1\) (one WC–uniform) and \(m=2\) (two WC–uniform) by BIC and record the selection frequency for \(m=2\).

We consider \(n\in \{1000,3000\}\) with \(R=6\) and \(R=4\) replicates, respectively. As in the no-covariate experiment, the smaller number of replicates at the larger sample size reflects the cost of repeated multi-start EM fitting. The covariate scale is not intrinsic to the model. If a scalar covariate is replaced by a one-to-one affine transformation \(X^*=a+bX\), \(b\ne 0\), then the same softmax weight functions can be represented by transformed regression coefficients, provided the basis is transformed accordingly. In particular, using a normalized covariate on [0, 1] is equivalent to fitting the same model after rescaling the original covariate range. Thus, the choice \(X\in [0,1]\) is a normalization for simulation design rather than a modeling restriction. Angle errors are wrapped to \([-\pi ,\pi )\). For the weight functions we report average bias over a 15-point grid \(x_g\in [0,1]\):

$$ \textrm{Bias}\big (\pi '_k(\cdot )\big )=\frac{1}{15}\sum _g \big (\widehat{\pi }'_k(x_g)-\pi '_k(x_g)\big ), $$

and compute a covariate-averaged Kullback–Leibler divergence between the true and fitted \(f(\cdot \mid x)\) on a fine \(\theta \)-grid and the same x-grid.

Results and interpretation. Table 5 shows the BIC hit-rate for the correct \(m=2\) rising from \(n=1000\) to \(n=3000\), with a concurrent decline in covariate-averaged KL risk. Shape errors in Table 6 remain small for \((\phi _1,\rho _1)\); \(\phi _2\) is noisier at \(n=1000\) because \(\rho _2\) is low, precisely as in the no-covariate study. Table 7 reveals mild overestimation of the uniform share (and corresponding underestimation of \(\pi '_1(\cdot )\) and \(\pi '_2(\cdot )\)) at these replicate counts; as in Sect. 4.1, this behavior should be interpreted as a near-boundary effect rather than only as an overlap effect. The true uniform weight function is small over the covariate range, so finite-sample estimation is sensitive to the lower boundary at zero and to weak separation between the uniform layer and low-concentration WC tails. Overall, the results confirm the intended separation: covariates modulate regime incidence through \(\pi '_k(\cdot )\) while global component shapes remain stable and interpretable.

Table 5 Covariate-dependent simulation: model selection (BIC) and covariate-averaged KL divergence for the weights-varying CD–MoKJ (two-component WC–uniform mixture) versus a one-component alternative (both with covariate-dependent weights)

Full size table

Table 6 Bias and RMSE for component shapes under CD–MoKJ (weights vary with x).

Full size table

Table 7 Average bias for \(\pi '_k(x)\) over a 15-point grid on [0, 1]

Full size table

5 Real-data application: wind direction mixtures

We analyze wind-direction records from two NOAA/NDBC stations using the standard meteorological archive. The goals are: (i) to fit unconditional WC–uniform mixtures that characterize baseline directional regimes and isolate diffuse mass; and (ii) to fit a covariate-dependent specification in which only the class probabilities vary with time of day and wind speed while component shapes remain global. Notation and estimation follow Sects. 23. For unconditional fits we report WC-core weights \(w_j=\pi '_j\) and the uniform weight \(\pi _u\); for covariate-dependent fits we retain global \((\phi _k,\rho _k)\) and model \(\pi '_1(x),\ldots ,\pi '_m(x)\) and \(\pi _u(x)\) via softmax regression.

5.1 Unconditional mixtures (no covariates)

Data and preprocessing. We consider two contrasting stations: 41047 (open ocean, NE Bahamas), typically reflecting persistent large-scale regimes, and MLRF1 (Molasses Reef, coastal Florida Keys), where sea/land-breeze cycles and coastline geometry induce broader, multimodal structure. Wind direction (WDIR) is parsed and converted to radians on \([0,2\pi )\), then thinned by retaining every sixth record to attenuate short-range dependence; models are subsequently fit under an independent-sampling working assumption. For 41047 we use July 2019. For MLRF1 the July slice is sparse, so we use the full year 2019—after standard cleaning—to stabilize estimation and uncertainty summaries.

Model and estimation. We fit a WC–uniform mixture

$$ f(\theta )\;=\;\sum _{j=1}^{K} w_j\, \textrm{WC}(\theta \mid \phi _j,\rho _j)\;+\;\pi _u\,\frac{1}{2\pi }, \qquad \sum _{j=1}^{K} w_j+\pi _u=1, $$

where \(w_j=\pi '_j\) is the unconditional weight of the jth WC core and \(\pi _u\) is the unconditional uniform-background weight. selecting \(K\in \{2,3\}\) by BIC. Maximum likelihood is computed by EM with responsibility-weighted first trigonometric moments for \((\phi _j,\rho _j)\) (cf. (10)). Uncertainty is summarized by nonparametric bootstrap percentile intervals (resampling timestamps and refitting at the BIC-selected K). For the uniform weight, these intervals are interpreted descriptively, because inference near \(\pi _u=0\) is affected by the boundary of the parameter space and standard percentile bootstrap intervals need not have their nominal regular coverage in such nonregular settings (Andrews 2000; Davison and Hinkley 1997). We therefore also report a bootstrap-calibrated likelihood-ratio p-value for testing \(H_0:\pi _u=0\) against the WC–uniform alternative, using the boundary test described in Sect. 3.6. Angles \(\phi _j\) are reported on \([0,2\pi )\); circular intervals unwrap locally around the point estimate.

Results and interpretation. Figure 1 overlays fitted densities and the uniform baseline \(\pi _u/(2\pi )\) on density-scaled circular histograms; the BIC-selected K is indicated. Tables 8 and 9 report point estimates and 95% bootstrap intervals for \((\phi _j,\rho _j,w_j)\) and \(\pi _u\).

Fig. 1

MoKJ (WC–uniform) EM fits with BIC-selected K: 41047 (top; July 2019) and MLRF1 (bottom; full 2019 due to sparse July). Dashed: fitted density; dotted: uniform baseline \(\pi _u/(2\pi )\)

Table 8 MoKJ (WC–uniform) parameter estimates with 95% bootstrap CIs for NDBC 41047.

Full size table

Table 9 MoKJ (WC–uniform) parameter estimates with 95% bootstrap CIs for NDBC MLRF1.

Full size table

For 41047, multiple concentrated regimes with small \(\pi _u\) are identified, consistent with persistent trade-wind flow; the large \(\rho _j\) indicate narrow spread. For MLRF1, broader multimodality and a larger \(\pi _u\) reflect diffuse directions during sea/land-breeze transitions. These patterns match the simulations: when background mass is negligible and a single sharp regime dominates, centers are tight; when transitional diffusion is present, the uniform layer absorbs it rather than forcing components to over-sharpen.

5.2 Covariate-dependent mixtures (weights vary with covariates)

Data and covariates. We use the same stations and periods. Directions are converted to radians on \([0,2\pi )\), and the covariate vector is

$$\begin{aligned} X=\bigl [1,\ \sin (2\pi \,\text {hour}/24),\ \cos (2\pi \,\text {hour}/24), \,\text {speed}_z\bigr ], \end{aligned}$$

where \(\text {speed}_z\) is standardized wind speed. To mitigate residual dependence, every sixth record is retained prior to fitting. This thinning rule retains one observation from each block of six consecutive records, which reduces the strongest short-range serial dependence while preserving sufficient sample size and coverage across the diurnal cycle. The resulting fits are therefore interpreted under an independent-sampling working likelihood, rather than as a full time-series model.

Model and estimation. We fit

$$ f(\theta \mid x)=\pi _u(x)\frac{1}{2\pi }+\sum _{k=1}^K \pi '_k(x)\,\textrm{WC}(\theta \mid \phi _k,\rho _k), \qquad \sum _{k=1}^K \pi '_k(x)+\pi _u(x)=1, $$

with class probabilities parameterized by a multinomial-logit link in X:

$$ \Pr (C=c\mid X=x)=\textrm{softmax}\bigl (B^\top X\bigr )_c,\qquad c=0,1,\dots ,K, $$

where \(c=0\) corresponds to the uniform background and \(c=1,\dots ,K\) to the WC components; an identifiability constraint is imposed as in Sect. 3. We choose \(K\in \{2,3\}\) by BIC and estimate via a covariate-dependent EM: E-step responsibilities from the current softmax; M-steps for \((\phi _k,\rho _k)\) using responsibility-weighted first moments; and B by maximizing \(\sum _i\sum _c r_{ic}\log \Pr (C=c\mid X_i)\).

Results. Figure 2 displays density overlays by hour bins at median speed. Figure 3 shows estimated weights versus hour at three speed levels (low/median/high), emphasizing the uniform fraction and the leading component. For both sites BIC selected \(K=3\). Global component parameters and representative weights (hours 0, 12, and 18 at median speed) are given in Tables 10, 11.

Fig. 2

CD–MoKJ overlays by hour bins at median speed: 41047 (top; July 2019) and MLRF1 (bottom; 2019 full year). Dashed: fitted \(f(\theta \mid x)\); dotted: uniform baseline \(\pi _u(x)/(2\pi )\)

Fig. 3

Weight effects versus hour at three speed levels (low/median/high). Top: uniform weight; bottom: leading component weight (Comp 1)

Table 10 CD–MoKJ for 41047 with covariates hour of day (periodic) and wind speed (standardized).

Full size table

Table 11 CD–MoKJ for MLRF1 with the same covariates as 41047.

Full size table

Interpretation and link to theory. For 41047, a dominant regime persists across hours with very small \(\pi _u(x)\) and high concentrations (\(\rho _k\approx 0.78\)–0.98). Allowing weights to depend on hour and speed adjusts regime incidence without moving sharp centers—exactly the separation targeted by the WC–uniform parameterization and reflected in the simulations. For MLRF1, diurnal modulation reallocates mass among regimes and elevates \(\pi _u(x)\) near transition hours, capturing sea/land-breeze dynamics; the explicit background buffers diffuse directions so WC components need not over-sharpen their peaks. These outcomes align with Sect. 4 and the identifiability/EM guarantees in Sect. 3.

6 Comparative study: CD–MoKJ vs. alternatives
6.1 Design and competitors

We benchmark the covariate-dependent mixture of Kato–Jones (CD–MoKJ), implemented via its WC–uniform representation with covariate-dependent weights, against four alternatives: (i) covariate-dependent von Mises with a uniform background (CD–MvM); (ii) MoKJ without covariates; (iii) von Mises with a uniform background without covariates (MvM); (iv) von Mises kernel density estimation (KDE–vM; bandwidth chosen by leave-one-out cross-validation). The two vM–uniform competitors are included specifically to separate the effect of adding an explicit uniform background from the effect of replacing symmetric vM kernels by the heavier-shouldered KJ/WC cores. The covariate vector matches Sect. 5,

$$ X=\bigl [1,\ \sin (2\pi \,\textrm{hour}/24),\ \cos (2\pi \,\textrm{hour}/24),\ \textrm{speed}_z\bigr ]. $$

For the comparative benchmark, the number of circular components \(K\in \{2,3\}\) is selected once from the unconditional fit and then held fixed across all competing methods. This choice gives a common component dimension for the cross-method comparison and avoids favoring one model class through a different model-selection step. However, this two-step strategy can be conservative: if covariates explain part of the regime incidence, the marginal mixture may merge regimes that become more clearly separated conditionally on X. For this reason, we also regard the fully covariate-dependent BIC over K as a natural sensitivity check, with parameter count adjusted for the additional multinomial-logit coefficients. In the present application, the selected values of K were stable over the small grid considered, and the substantive conclusions were unchanged. Predictive performance is evaluated with fivefold time-blocked cross-validated log-likelihood (CVLL), using contiguous folds to limit residual serial dependence.

6.2 Diagnostic profile and theoretical expectations

We summarize descriptors that indicate which family should perform best for a given site/month: (i) resultant length \(\bar{R}=\frac{1}{n}\bigl \Vert \sum _{j=1}^n e^{\textrm{i}\theta _j}\bigr \Vert \) (overall concentration); (ii) Rayleigh’s uniformity p-value; (iii) circular skewness and excess kurtosis; (iv) an estimate of the average uniform share \(\overline{\pi _u(x)}\) from a quick CD–MoKJ fit with \(K\in \{2,3\}\) chosen by BIC. When \(\bar{R}\) is large, \(\overline{\pi _u(x)}\approx 0\), and skewness is mild, von Mises components efficiently approximate the modal core and tend to dominate in CVLL. With multiple regimes, skewed shoulders, and nonnegligible diffuse mass, the heavier-tailed WC cores plus explicit uniform background of CD–MoKJ are favored—consistent with Sect. 4.

6.3 Station 41047 (July 2019; open ocean)

The July record is nearly unimodal and sharply concentrated, with negligible diffuse mass. Table 12 reports CVLL: CD–MvM achieves the highest value; KDE–vM and static MvM are close; WC-based specifications lag behind. Figure 4 shows CD–MvM overlays by hour at median speed; the estimated uniform level is essentially zero across the day. These outcomes match the diagnostics (large \(\bar{R}\), tiny \(\overline{\pi _u(x)}\)) and the inductive bias of von Mises components in single tight-mode settings.

Table 12 Model comparison on 41047 (July 2019).

Full size table

Fig. 4

Station 41047 (July 2019). CD–MvM overlays by hour at median speed. Dashed: fitted density; dotted: uniform level \(\pi _u(x)/(2\pi )\). The background remains negligible

6.4 Station MLRF1 (2019; coastal)

The coastal record shows multiple regimes and pronounced diurnal modulation, with broader spread at transitions. Figure 5 displays CD–MoKJ overlays by hour at median speed: incidence shifts among regimes over the cycle and the uniform share peaks at transition hours. Under the same time-blocked protocol, the numerical CVLL results are reported in Table 13. In this comparison, KDE–vM gives the highest CVLL, with CD–MvM and static MvM close behind, while CD–MoKJ is less competitive by this purely predictive likelihood criterion. The pattern nevertheless mirrors the diagnostic interpretation (moderate \(\bar{R}\), nontrivial \(\overline{\pi _u(x)}\), asymmetry): WC tails avoid over-sharpening in overlap regions, and the explicit uniform layer isolates genuinely diffuse directions.

Table 13 Model comparison on MLRF1 (2019).

Full size table

Fig. 5

Station MLRF1 (2019). CD–MoKJ overlays by hour at median speed. Mass shifts among regimes over the diurnal cycle, and the uniform share increases at transition hours, consistent with the model’s separation of diffuse and structured components

6.5 Synthesis and link to theory

Across sites, the empirical ordering shows that predictive likelihood and interpretability need not select the same model. For 41047, the record is close to unimodal and highly concentrated, and the covariate-dependent von Mises mixture achieves the highest fivefold time-blocked CVLL. This is consistent with the theoretical expectation that symmetric vM components are efficient when directional regimes are sharp and the diffuse background is weak. For MLRF1, the coastal record exhibits multiple regimes, stronger diurnal modulation, and nonnegligible diffuse transitional mass. In that setting, KDE–vM and the unconditional MoKJ fit achieve the highest CVLL values, while CD–MoKJ remains useful for interpreting covariate-dependent regime incidence and the uniform background. Thus, the comparison should be read as complementary: CVLL measures out-of-sample predictive density, whereas the WC–uniform structure provides regime-level diagnostics for skewness, diffuse mass, and transition behavior. In both stations, K is fixed by BIC prior to adding covariates, so the comparisons are made at a common component dimension rather than by allowing each method to choose a different complexity.

Table 14 Numerical comparison of fivefold time-blocked CV log-likelihoods across the two stations.

Full size table

7 Conclusion

This work develops a covariate-dependent mixture of Kato–Jones submodels with an explicit uniform background (CD–MoKJ) for circular data exhibiting multimodality, skewness, and diffuse transitional mass. A boundary–concentration reparameterization reveals each KJ component as a convex combination of a Wrapped–Cauchy core and a uniform layer (WC–uniform), separating regime shape from genuinely unstructured directions and yielding an identifiable, numerically stable specification.

On the theory side, identifiability for the covariate-dependent model (up to label permutation) follows from trigonometric-moment arguments under mild ordering and compactness. The presence of a uniform background is testable via a likelihood-ratio statistic with a chi-bar-square limit, with practical calibration by parametric bootstrap. Estimation relies on a finite-dimensional sieve–EM: a strictly concave multinomial-logit update governs covariate-dependent class probabilities, while WC parameters admit closed-form moment updates. Standard EM guarantees ensure monotone ascent and stationarity of limit points, and a fixed-dimension BIC consistently selects the number of circular components.

Simulation studies reinforce these properties. Without covariates, bias and RMSE diminish with n and a generalized MSE contracts rapidly, reflecting identifiability and stable updates. With covariate-dependent weights, predictive performance improves while component centers remain fixed, matching the intended separation between regime incidence and regime shape. Applications to NOAA wind directions illustrate the model’s operating regime. For an open-ocean month dominated by a single sharp mode and negligible diffuse mass, a covariate-dependent von Mises mixture attains the best out-of-sample likelihood—unsurprising when heavier tails and a background layer offer little gain. For a coastal record with multiple regimes and pronounced transition periods, the WC–uniform formulation provides interpretable diagnostics: the uniform layer isolates diffuse directions, and WC cores avoid over-sharpening in overlap regions, although the best predictive likelihood in this case is achieved by more flexible kernel-based or von Mises competitors. The present formulation deliberately lets covariates affect the mixture weights while keeping the component shapes fixed. This restriction is appropriate when predictors mainly alter the incidence of directional regimes, as in the wind-direction applications considered here, but it may be limiting in settings where the mean direction, concentration, or skewness of a regime changes with environmental conditions. A direct extension is to allow selected shape parameters, such as \(\mu _k\), \(\lambda _k\), or \(\rho _k\), to depend smoothly on covariates through low-dimensional link functions or penalized sieves, while preserving ordering and compactness constraints needed for identifiability and stable EM fitting. Other natural extensions include introducing temporal dependence, for example through hidden Markov or autoregressive circular mixtures with the uniform layer acting as a transient state; using richer sieves or penalties for higher-dimensional covariates; developing Bayesian formulations that encode constraints while retaining the core–plus–background structure; and spatial pooling across stations with robust variance adjustments for residual dependence.

In sum, CD–MoKJ provides a principled and practical framework for circular data with regime structure and diffuse mass, combining interpretable components with covariate-responsive weights and delivering predictive gains precisely where classical von Mises mixtures are too rigid or prone to over-sharpening.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Modelling non-stationary extremal dependence through a geometric approach05.8324-07-2026
2Novel fast and asymptotically efficient estimation in weighted exponential family06.0622-07-2026
3Emulated FPE Order Identification for Autoregressive Models of Big Multivariate Time Series07.722-07-2026
4Group LASSO for multiple change-point detection in a generalized integer-valued autoregressive model09.1824-07-2026
5Convergence of systematic-scan and random-scan Gibbs samplers for bivariate discrete conditional distributions07.5321-07-2026
6Quantile adaptive feature screening for ultra-high dimensional longitudinal heterogeneous data08.9824-07-2026
7Best Fixed and Sequential Design for Bayesian Estimation of a Sum of Two Failure Rates of Exponential Distributions06.7920-07-2026
8Objective Bayesian Analysis for the Generalized Logistic Distribution04.7422-07-2026
9The Ensemble Learning to Determine Optimal Tuning Parameter of the Generalized Lasso in Spatial Clustering Analysis05.122-07-2026
10Estimation in a Hazard Regression Model with Smooth Transitions over Unknown Regimes03.7220-07-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 8.02. Источник: link.springer.com.