Phi coefficient

In statistics, the phi coefficient (also referred to as the "mean square contingency coefficient" and denoted by φ or r_φ) is a measure of association for two binary variables introduced by Karl Pearson^[1]. This measure is similar to the Pearson correlation coefficient in its interpretation. In fact, a Pearson correlation coefficient estimated for two binary variables will return the phi coefficient.^[2] The square of the Phi coefficient is related to the chi-squared statistic for a 2×2 contingency table (see Pearson's chi-squared test)^[3]

$\phi^2 = \frac{\chi^2}{n}$

where n is the total number of observations. Two binary variables are considered positively associated if most of the data falls along the diagonal cells. In contrast, two binary variables are considered negatively associated if most of the data falls off the diagonal. If we have a 2×2 table for two random variables x and y

	y = 1	y = 0	total
x = 1	$n 11$	$n 10$	$n_{1\bullet}$
x = 0	$n 01$	$n 00$	$n_{0\bullet}$
total	$n_{\bullet1}$	$n_{\bullet0}$	$n$

where n₁₁, n₁₀, n₀₁, n₀₀, are non-negative "cell cell counts" that sum to n, the total number of observations. The phi coefficient that describes the association of x and y is

$\phi = \frac{n_{11}n_{00}-n_{10}n_{01}}{\sqrt{n_{1\bullet}n_{0\bullet}n_{\bullet0}n_{\bullet1}}}$

Phi is related to the point-biserial correlation coefficient and Cohen's d and estimates the extent of the relationship between two variables (2 x 2).^[4]

Maximum values

Although computationally the Pearson correlation coefficient reduces to the phi coefficient in the 2×2 case, the interpretation of a Pearson correlation coefficient and phi coefficient must be taken cautiously. The Pearson correlation coefficient ranges from −1 to +1, where ±1 indicates perfect agreement or disagreement, and 0 indicates no relationship. The phi coefficient has a maximum value that is determined by the distribution of the two variables. If both have a 50/50 split, values of phi will range from −1 to +1. See Davenport El-Sanhury (1991) ^[5] for a thorough discussion.

References

^ Cramer, H. 1946. Mathematical Methods of Statistics. Princeton: Princeton University Press, p282 (second paragraph). ISBN 0691080046
^ Guilford, J. (1936). Psychometric Methods. New York: McGraw–Hill Book Company, Inc.
^ Everitt B.S. (2002) The Cambridge Dictionary of Statistics, CUP. ISBN 0-521-81099-x
^ Aaron, B., Kromrey, J. D., & Ferron, J. M. (1998, November). Equating r-based and d-based effect-size indices: Problems with a commonly recommended formula. Paper presented at the annual meeting of the Florida Educational Research Association, Orlando, FL. (ERIC Document Reproduction Service No. ED433353)
^ Davenport, E., & El-Sanhury, N. (1991). Phi/Phimax: Review and Synthesis. Educational and Psychological Measurement, 51, 821–828.

Statistics

Descriptive statistics

Continuous data

Location	Mean (Arithmetic, Geometric, Harmonic) · Median · Mode

Dispersion	Range · Standard deviation · Coefficient of variation · Percentile · Interquartile range

Shape	Variance · Skewness · Kurtosis · Moments · L-moments

Count data

Index of dispersion

Summary tables

Grouped data · Frequency distribution · Contingency table

Dependence

Pearson product-moment correlation · Rank correlation (Spearman's rho, Kendall's tau) · Partial correlation · Scatter plot

Statistical graphics

Bar chart · Biplot · Box plot · Control chart · Correlogram · Forest plot · Histogram · Q-Q plot · Run chart · Scatter plot · Stemplot · Radar chart

Data collection

Designing studies	Effect size · Standard error · Statistical power · Sample size determination

Survey methodology	Sampling · Stratified sampling · Opinion poll · Questionnaire

Controlled experiment	Design of experiments · Factorial experiment · Randomized experiment · Random assignment · Replication · Blocking · Optimal design

Uncontrolled studies	Natural experiment · Quasi-experiment · Observational study

Statistical inference

Statistical theory	Sampling distribution · Sufficient statistic · Meta-analysis

Bayesian inference	Bayesian probability · Prior · Posterior · Credible interval · Bayes factor · Bayesian estimator · Maximum posterior estimator

Frequentist inference	Confidence interval · Hypothesis testing · Likelihood-ratio

Specific tests	Z-test (normal) · Student's t-test · F-test · Pearson's chi-squared test · Wald test · Mann–Whitney U · Shapiro–Wilk · Signed-rank · Kolmogorov–Smirnov test

General estimation	Mean-unbiased · Median-unbiased · Maximum likelihood · Method of moments · Minimum distance · Density estimation

Correlation and regression analysis

Correlation	Pearson product-moment correlation · Partial correlation · Confounding variable · Coefficient of determination

Regression analysis	Errors and residuals · Regression model validation · Mixed effects models · Simultaneous equations models

Linear regression	Simple linear regression · Ordinary least squares · General linear model · Bayesian regression

Non-standard predictors	Nonlinear regression · Nonparametric · Semiparametric · Isotonic · Robust

Generalized linear model	Exponential families · Logistic (Bernoulli) · Binomial · Poisson

Partition of variance	Analysis of variance (ANOVA) · Analysis of covariance · Multivariate ANOVA · Degrees of freedom

Categorical, multivariate, time-series, or survival analysis

Categorical data	Cohen's kappa · Contingency table · Graphical model · Log-linear model · McNemar's test

Multivariate statistics	Multivariate regression · Principal components · Factor analysis · Cluster analysis · Copulas

Time series analysis	Decomposition (Trend · Stationary process) · ARMA model · ARIMA model · Vector autoregression · Spectral density estimation

Survival analysis	Survival function · Kaplan–Meier · Logrank test · Failure rate · Proportional hazards models · Accelerated failure time model

Applications

Biostatistics	Bioinformatics · Biometrics · Clinical trials & studies · Epidemiology · Medical statistics · Pharmaceutical statistics

Engineering statistics	Methods engineering · Probabilistic design · Process & Quality control · Reliability · System identification

Social statistics	Actuarial science · Census · Crime statistics · Demography · Econometrics · National accounts · Official statistics · Population · Psychometrics

Spatial statistics	Cartography · Environmental statistics · Geographic information system · Geostatistics · Kriging

Category · Portal · Outline · Index

Categories:

Categorical data
Statistical dependence
Statistical ratios
Summary statistics for contingency tables

Wikimedia Foundation. 2010.

Игры ⚽ Поможем сделать НИР

Look at other dictionaries:

phi coefficient — noun an index of the relation between any two sets of scores that can both be represented on ordered binary dimensions (e.g., male female) • Syn: ↑phi correlation, ↑fourfold point correlation • Topics: ↑statistics • Hypernyms: ↑ … Useful english dictionary
phi correlation — noun an index of the relation between any two sets of scores that can both be represented on ordered binary dimensions (e.g., male female) • Syn: ↑phi coefficient, ↑fourfold point correlation • Topics: ↑statistics • Hypernyms: ↑ … Useful english dictionary
Phi (letter) — Phi (uppercase Φ, lowercase φ or Unicode|ϕ), pronounced [IPA|fī] in modern Greek and as [IPA|faɪ] in English, is the 21st letter of the Greek alphabet. In modern Greek, it represents [IPA|f] , a voiceless labiodental fricative. In Ancient Greek… … Wikipedia
Coefficient de marée — Calcul de marée Le calcul de marée est la méthode utilisée en navigation maritime pour estimer la hauteur d eau, dans un lieu et à un instant donné, en prenant en compte l influence de la marée. Les marées sont le résultat de l attraction de la… … Wikipédia en Français
Matthews correlation coefficient — The Matthews correlation coefficient is used in machine learning as a measure of the quality of binary (two class) classifications. It takes into account true and false positives and negatives and is generally regarded as a balanced measure which … Wikipedia
Activity coefficient — An activity coefficient [ [http://www.iupac.org/goldbook/A00116.pdf Gold Book definition] ] is a factor used in thermodynamics to account for deviations from ideal behaviour in a mixture of chemical substances. In an ideal mixture the… … Wikipedia
Effect size — In statistics, an effect size is a measure of the strength of the relationship between two variables in a statistical population, or a sample based estimate of that quantity. An effect size calculated from data is a descriptive statistic that… … Wikipedia
Contingency table — In statistics, a contingency table (also referred to as cross tabulation or cross tab) is a type of table in a matrix format that displays the (multivariate) frequency distribution of the variables. It is often used to record and analyze the… … Wikipedia
Cramér's V — Cramér s V (φc) In statistics, Cramér s V (sometimes referred to as Cramér s phi and denoted as φc) is a popular[citation needed] measure of association between two nominal variables, giving a value between 0 and +1 (inclusive). It… … Wikipedia
distribution free statistic — noun a statistic computed without knowledge of the form or the parameters of the distribution from which observations are drawn • Syn: ↑nonparametric statistic • Topics: ↑statistics • Hypernyms: ↑statistic • Hyponyms: ↑ … Useful english dictionary

Academic Dictionaries and Encyclopedias

Phi coefficient

Maximum values

See also

References

Look at other dictionaries:

Share the article and excerpts

Academic Dictionaries and Encyclopedias

Wikipedia

Phi coefficient

Maximum values

See also

References

Look at other dictionaries:

Share the article and excerpts

Direct link