This may take 2 months or more, since drafts are reviewed in no specific order. There are 2,468 pending submissions waiting for review.
Review waiting, please be patient.
This may take 2 months or more, since drafts are reviewed in no specific order. There are 2,468 pending submissions waiting for review.
Where to get help
How to improve a draft
You can also browse Wikipedia:Featured articles and Wikipedia:Good articles to find examples of Wikipedia's best writing on topics similar to your proposed article. Improving your odds of a speedy review To improve your odds of a faster review, tag your draft with relevant WikiProject tags using the button below. This will let reviewers know a new draft has been submitted in their area of interest. For instance, if you wrote about a female astronomer, you would want to add the Biography, Astronomy, and Women scientists tags. Editor resources
Reviewer tools
|
Invariant coordinate selection (ICS) is a multivariate statistical technique used to identify interesting structures in high-dimensional data.[1] It is commonly applied in outlier detection,[2] independent component analysis (ICA),[3] and robust dimension reduction. ICS generalizes principal component analysis (PCA) by simultaneously diagonalizing two scatter matrices, typically the sample covariance matrix and a higher-order or robust scatter matrix.[1]
Unlike PCA, which depends on the scale of the variables and captures only second-order structure, ICS is invariant under full-rank affine transformations of the data and can reveal non-Gaussian features of the underlying distribution.[1] The core idea is to find a linear transformation that simultaneously diagonalizes two scatter matrices, yielding coordinates whose ordering reflects the disagreement between the two scatter measures. The method draws heavily on the robust statistics literature for its repertoire of scatter functionals.[3]
The simultaneous use of two scatter matrices for data transformation predates ICS and originated in the clustering literature. Art, Gnanadesikan and Kettenring (1982) introduced a data-based metric for cluster analysis in which an estimate of the within-cluster covariance matrix is computed iteratively without knowledge of the cluster labels, and the data are then rescaled by its inverse square root.[4] This procedure, implemented as PROC ACECLUS in SAS, is generally regarded as the earliest instance of the simultaneous-diagonalization idea later generalized by ICS.[1] Caussinus and Ruiz-Gazen (1990) developed a closely related generalized principal component analysis for projection pursuit and cluster identification.[5]
In parallel, the same algebraic structure was being investigated in signal processing under the headings of independent component analysis and blind source separation. Cardoso's Fourth-Order Blind Identification (FOBI) jointly diagonalizes the covariance matrix and a fourth-moment scatter matrix to recover independent sources.[6] Oja and collaborators subsequently showed that any pair of affine-equivariant scatter matrices possessing the so-called independence property can be used to recover an independent component model, placing FOBI and related procedures in a common statistical framework.[7][8]
A third influence comes from the robust statistics community, where most of the alternative scatter functionals later adopted by ICS were originally introduced—including the M-estimators of multivariate scatter studied by Maronna,[9] the distribution-free M-estimator of Tyler,[10] the symmetrized M-estimator of Dümbgen,[11] and the minimum covariance determinant (MCD) estimator of Rousseeuw.[12]
These threads were unified by Tyler, Critchley, Dümbgen and Oja (2009), who introduced the name invariant coordinate selection, characterized its affine-invariance properties, and established it as a general-purpose exploratory method that subsumes the earlier proposals as special cases.[1]
Let be a sample of observations in , with sample mean .
A scatter matrix is a symmetric positive-definite matrix-valued functional of the data that is affine equivariant, meaning that for any non-singular matrix and vector ,
The sample covariance matrix is the canonical example, but many alternatives drawn from the robust statistics literature are available (see below).[1]
Given two scatter matrices and , ICS solves the generalized eigenvalue problem
producing eigenvalues and a matrix of corresponding eigenvectors normalized so that .[1] The invariant coordinates are the columns of
The eigenvalues quantify how strongly the two scatter matrices disagree along each direction. For any unit vector , the ratio can be interpreted as a generalized kurtosis measure along that direction, so the invariant coordinates are ordered by a notion of multivariate kurtosis.[1][7] Coordinates associated with extreme eigenvalues are those typically of greatest interest for outlier detection or clustering.[2]
When is taken to be the identity matrix and the sample covariance matrix, ICS reduces to ordinary principal component analysis; PCA is therefore a special case in which only second-order structure is exploited.[1]
The choice of the pair determines which features ICS will emphasize. Common scatter functionals used in practice include:
Pairing two functionals with contrasting sensitivities—for instance the covariance matrix with a robust scatter, or with the fourth-moment scatter—produces invariant coordinates that highlight different features of the data, such as outliers, clusters, or non-Gaussian sources.[3]
A scatter matrix generalizes the covariance matrix as a summary of multivariate spread, and different scatter matrices emphasize different aspects of the data distribution. ICS typically pairs the ordinary covariance matrix with a second scatter matrix that is either robust to outliers or sensitive to higher-order moments such as kurtosis.[3]
For elliptically symmetric distributions, such as the multivariate normal distribution, any two affine-equivariant scatter matrices are proportional, so all generalized eigenvalues coincide and no direction is preferred.[1] Departures from this equality therefore signal departures from ellipticity—such as the presence of clusters, skewness, or outliers—and the corresponding invariant coordinates highlight the directions in which these features are most pronounced.
{{cite journal}}: CS1 maint: DOI inactive as of May 2026 (link)
Category:Multivariate statistics Category:Matrix decompositions Category:Dimension reduction Category:Robust statistics Category:Exploratory data analysis Category:Cluster analysis Category:Statistical methods
Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.