Reference
Definitions, notation, assumptions and common errors, for looking up something you have already met. Each entry carries what it connects to. For a first encounter, start from Learn; to see the whole structure at once, open the map.
Analysis
- First-Order Differential Equations
An equation relating a function to its own derivative, the two methods that solve the first-order cases, separation of variables and the integrating factor, and the initial condition that selects one solution from the infinite family.
6 connections
- Functions, Domains and Inequalities
What a function is as a rule with a stated domain, how the domain is found from the operations that would fail without it and written in interval notation, and separately how the inequalities and exponential equations describing such sets are solved, including the domain check that discards roots the algebra invents.
7 connections
Related: Trigonometry
Used by: Limits and Continuity, The Derivative
Related: Concept Hierarchies as Ordered Structures
Required by: Sequences and Their Limits, The Natural Logarithm
and 1 more
Valuing a Stream of Dated Payments - Numerical Solution of Differential Equations
Stepping methods that approximate a solution when no formula exists, carried out by hand for Euler and improved Euler, and separately judged: the order of accuracy that says how each responds to a smaller step, the stability limit that can make a smaller step necessary rather than merely better, and what a computed answer does and does not establish.
5 connections
- Second-Order Linear Differential Equations
Constant-coefficient equations of the form
, the characteristic polynomial that substitutingproduces, and the three shapes of general solution its roots determine. Two exponentials, an exponential with an factor, or a decaying oscillation. 5 connections
Related: Linear Transformations, Vector Spaces and Subspaces
Requires: Complex Numbers, First-Order Differential Equations
Required by: Systems of Linear Differential Equations
- Sequences and Their Limits
Infinite lists of numbers indexed by
, what it means for one to converge, the – definition that makes the claim checkable, and the monotone convergence theorem that establishes a limit exists without naming it. 5 connections
Related: Numerical Solution of Differential Equations
Requires: Functions, Domains and Inequalities, Limits and Continuity
Used by: Convergence of Infinite Series
Required by: Valuing a Stream of Dated Payments
- Systems of Linear Differential Equations
Coupled equations written as
, solved by the eigenvalues and eigenvectors of: each eigenpair contributes a solution , and the eigenvalues' signs and imaginary parts decide whether trajectories decay, grow, or circulate. 5 connections
Related: Diagonalization
Requires: Complex Numbers, Eigenvalues and Eigenvectors
- The Natural Logarithm
The function
, the product law that follows from that area rather than being assumed, the numberdefined as the point where the area reaches 1, and the exponential recovered as its inverse. 7 connections
- Trigonometry
Sine and cosine defined as coordinates on the unit circle, the exact values for the special angles obtained from triangle geometry, the identities derived from the circle's equation and the angle-addition formulas, and the equations whose solutions repeat forever.
5 connections
Related: Complex Numbers
Used by: Limits and Continuity, Orthogonality and Projection
and 1 more
The DerivativeRelated: Functions, Domains and Inequalities
Behavioural Economics
- Choices No Utility Function Represents
Expected utility as a representation theorem rather than a prediction, the common-consequence pair whose two preferences impose contradictory demands on the same utility number, what the independence axiom asserts and where it fails, and the descriptive model built to accommodate the failure.
4 connections
Analogous to: Validity as an Argument for an Interpretation
Contrasts with: Constrained Optimization
Related: Valuing a Stream of Dated Payments
Calculus
- Conservative Fields, Potentials and Path Independence
When a line integral depends only on its endpoints: the potential whose gradient is the field, the cross-partial test that rules a potential out in one direction and establishes one only on a simply connected region, and the punctured plane where the test passes and no potential exists.
- Convergence of Infinite Series
An infinite sum defined as the limit of its partial sums, the tests that decide whether that limit exists, and the distinction between absolute and conditional convergence that decides whether the terms may be reordered.
6 connections
Related: Techniques of Integration
Requires: Limits and Continuity
Used by: Power Series, Taylor Expansion and the Remainder
Related: Numerical Solution of Differential Equations
Required by: Power Series, Taylor Expansion and the Remainder
- Double Integrals and Fubini's Theorem
The double integral as the Riemann construction with area elements in place of widths, Fubini's theorem as what reduces that two-dimensional limit to two ordinary integrals, and why the freedom to choose an order is a property of the rectangle rather than of integration.
3 connections
Requires: The Definite Integral
Required by: Green's Theorem
- Green's Theorem
Trading a closed boundary integral for a double integral over the region it encloses: the three hypotheses each doing work, the choice of field that turns the region integrand into a constant and computes an area from its boundary alone, and the enclosed singularity that voids the theorem.
4 connections
- L'Hôpital's Rule
A theorem that resolves
and by differentiating numerator and denominator separately, the conditions that make it apply, the other indeterminate forms that must be rewritten as quotients first, and the cases where it gives a wrong answer or no answer at all. 2 connections
Requires: Limits and Continuity, The Derivative
- Limits and Continuity
What it means for a function to approach a value, made precise enough to carry the weight the derivative and the integral put on it. The limit laws, the indeterminate forms they cannot settle, and continuity as the condition that lets a limit be read off by substitution.
8 connections
Used by: The Definite Integral, The Derivative
Required by: Convergence of Infinite Series, L'Hôpital's Rule
- Line Integrals and Path Parametrisation
Integrating a vector field along a route rather than over an interval: how a parametrisation turns the line integral into an ordinary single-variable integral, why the value is a property of the oriented curve, and why the answer generally depends on which route was taken.
5 connections
- Partial Derivatives, the Gradient and Critical Points
Partial derivatives as rates along the coordinate axes, the gradient that assembles them into a vector determining the rate in every other direction, and the Hessian that separates a minimum from a maximum from the saddle that one variable cannot produce.
5 connections
Related: Quadratic Forms and Definiteness
Requires: Cross Products and Geometry in Space, The Derivative
Used by: Double Integrals and Fubini's Theorem
Required by: Conservative Fields, Potentials and Path Independence
- Power Series, Taylor Expansion and the Remainder
A power series as a representation of a function on an interval, the Taylor coefficients built from derivatives at a point, and the Lagrange remainder that decides both how far a truncation can be trusted and whether the series represents the function at all.
4 connections
Requires: Convergence of Infinite Series, The Derivative
Related: The Natural Logarithm
- Techniques of Integration
Integration by parts, which inverts the product rule; partial fractions, which splits a rational integrand into terms the power and logarithm rules reach; and improper integrals, where an unbounded interval or an unbounded integrand turns the integral into a limit that may or may not converge.
6 connections
Requires: Limits and Continuity, The Definite Integral
and 1 more
The DerivativeRelated: Convergence of Infinite Series
Required by: First-Order Differential Equations
Uses: The Natural Logarithm
- The Definite Integral
A limit of Riemann sums, read as signed area and as accumulated change. The fundamental theorem, which makes the integral computable by antidifferentiation instead of by summing, and the substitution rule that inverts the chain rule.
8 connections
Related: Linear Transformations
Requires: The Derivative
Required by: Double Integrals and Fubini's Theorem, First-Order Differential Equations
Uses: Limits and Continuity
- The Derivative
The limit of difference quotients, read as an instantaneous rate and as the slope of the best linear approximation. The rules for powers, products, quotients and compositions, each a consequence of that limit rather than a convention, and the three ways the limit can fail to exist.
11 connections
Category Theory
- Mappings That Preserve Structure
A category as objects, arrows and a composition satisfying two laws; a functor as a mapping that preserves those laws rather than merely matching up objects; the composite that detects a mapping which looks structural and is not; and naturality as a square that has to commute at every arrow.
3 connections
Analogous to: Linear Transformations
Related: Concept Hierarchies as Ordered Structures
Requires: Sets and Fields
Data Visualization and Reporting
- Choosing What the Reader Will Judge
A display turns numbers into a perceptual task, and the task decides how accurately the reader can answer. What the eye judges well and badly, why the scale and the baseline are claims rather than formatting, and what any chosen display conceals as the price of what it shows.
4 connections
Econometrics
- When a Coefficient Is Not an Effect
Why a least-squares coefficient estimated from observational data need not be the causal effect, the three mechanisms that break the exogeneity assumption, the formula that gives the size and direction of omitted-variable bias, and what an instrumental variable would have to satisfy to repair it.
4 connections
Contrasts with: Estimating Out-of-Sample Error
Elaborates: Linear Regression for Experimental Research
Requires: Linear Regression for Experimental Research, Unconfoundedness and Overlap
Experimental Research Design
- ANOVA for Experimental Research
One-way analysis of variance asks whether several group means can be treated as equal, by comparing the variation between groups with the variation within them. The decomposition is exact and the test is a ratio of mean squares. What rejection establishes is narrow, that the means are not all equal, and answering which ones differ is a separate question requiring planned contrasts or multiplicity-adjusted comparisons.
5 connections
Contrasts with: Hypothesis Tests for Experimental Research
Requires: Hypothesis Tests for Experimental Research
Contrasts with: Indicator Variables and Interactions, Testing Counts Against a Claim
Suggested after: Hypothesis Tests for Experimental Research
- Binary Outcome Models for Experimental Research
When the outcome is zero or one, its conditional mean is a probability, and two models compete to describe it. The linear probability model reports differences in probability directly and can predict outside the zero-to-one range; logistic regression keeps probabilities in range and reports odds ratios, which are neither risk ratios nor probability differences. The two differ in the functional form assumed for
, in whether fitted values stay inside, in the effect heterogeneity the link implies, in estimation, and in behaviour under extrapolation. The reporting scale is one consequence of that choice rather than the whole of it, and whichever scale is used must be stated. 3 connections
Contrasts with: Propensity Scores
Requires: Linear Regression for Experimental Research
Contrasts with: Linear Discriminants, and the Covariance They Assume
- Blocked and Paired Randomized Experiments
Grouping similar units before assignment, then randomizing within each group, reduces variance by removing baseline variation the comparison would otherwise carry. What it costs is a commitment: the analysis must follow the design. A blocked estimator weights stratum effects by stratum size, a paired estimator works on within-pair differences, and treating either as a free split across all units discards the precision the design was built to gain.
5 connections
Contrasts with: Randomized Assignment
Requires: Neyman Repeated-Sampling Inference, Randomized Assignment
Contrasts with: Matching for Causal Inference
Suggested after: Neyman Repeated-Sampling Inference
- Conditional Probability, Total Probability and Bayes' Rule
Conditional probability as a change of denominator rather than a change to the world, the asymmetry between
and that follows from it, and the decomposition of a probability over a partition that turns into Bayes' rule and makes the base rate govern a posterior as much as the likelihood does.4 connections
- Confidence Intervals for Experimental Research
An interval estimate is the same three ingredients as a test, rearranged: an estimate, a standard error, and a critical value that sets the coverage. What the confidence statement describes is the procedure's long-run behaviour, not the probability that one computed interval contains the parameter, and keeping that straight is what separates a reportable interval from a misreported one.
3 connections
Contrasts with: Hypothesis Tests for Experimental Research
Requires: Sampling Distributions and Standard Error
Best taken after: Estimators and How They Are Judged
- Covariance, Independence and the Variance of a Sum
Covariance as the cross term that decides whether the variance of a sum is the sum of the variances, independence as the stronger condition that the joint distribution factors, and the one-way implication between them that makes a zero correlation weaker evidence than it appears.
- Fisher Randomization Inference
Under the hypothesis that treatment changed nothing for anyone, every missing potential outcome is known: it equals the observed one. That makes the whole experiment recomputable under every allocation the design permitted, and the observed statistic can be placed in the list of values it could have taken. The result is an exact p-value that assumes no model, no distribution and no large sample, only the assignment mechanism.
4 connections
Contrasts with: Neyman Repeated-Sampling Inference
Requires: Potential Outcomes and the Fundamental Problem, Randomized Assignment
Contrasts with: Hypothesis Tests for Experimental Research
- Hypothesis Tests for Experimental Research
Every test in this family is one template: an estimate, a null value, and a standard error, assembled into a standardized distance and referred to a distribution that says how unusual such a distance would be if the null were true. Choosing the right test is choosing the right standard error and reference distribution for the design that produced the data, and what a large p-value licenses is far less than it is usually asked to carry.
9 connections
Contrasts with: Fisher Randomization Inference
Requires: Sampling Distributions and Standard Error
Suggested next: ANOVA for Experimental Research
Contrasts with: ANOVA for Experimental Research, Confidence Intervals for Experimental Research
Required by: ANOVA for Experimental Research, Linear Regression for Experimental Research
and 1 more
Testing Counts Against a ClaimSuggested after: Sampling Distributions and Standard Error
- Indicator Variables and Interactions
An indicator variable turns a category into a number a regression can use, and its coefficient is a difference in conditional means. An interaction lets a relationship differ by group, which changes what every other coefficient in the model means: once an interaction is present, the main effect is the effect at the reference value, not an overall effect.
4 connections
Contrasts with: ANOVA for Experimental Research
Requires: Linear Regression for Experimental Research
Used by: Regression Adjustment in Experiments
Suggested after: Linear Regression for Experimental Research
- Inverse Probability Weighting
Weighting each unit by the inverse of its probability of receiving the treatment it got targets a pseudo-population in which assignment is unrelated to the measured covariates. Under the true propensity score and the identification assumptions that is a population-level result; in practice the weights are estimated, so the balance actually achieved in the weighted sample is something to check rather than assume. The arithmetic is simple and its failure mode is specific: when some units had almost no chance of their observed treatment, their weights explode, and an estimate that rests on a handful of heavily weighted observations is reporting an overlap problem rather than an effect.
4 connections
Contrasts with: Matching for Causal Inference
Requires: Propensity Scores, Unconfoundedness and Overlap
Suggested after: Propensity Scores
- Linear Regression for Experimental Research
Least squares fits a line, or a hyperplane, by minimising squared residuals, and supplies coefficients with standard errors from which tests and intervals follow. Two things it does not supply: a causal reading, which comes from the design rather than the fit, and evidence of a good model, which R-squared does not measure. A randomized treatment effect can be entirely credible with a low R-squared.
12 connections
Requires: Hypothesis Tests for Experimental Research
Suggested next: Indicator Variables and Interactions
Used by: Regression Adjustment in Experiments
Elaborated by: When a Coefficient Is Not an Effect
Related: Orthogonality and Projection
Required by: Binary Outcome Models for Experimental Research, Choosing a Regression Form
- Matching for Causal Inference
Matching builds a comparison by pairing each unit with a similar unit under the opposite treatment, then comparing outcomes. Its design choices, the distance, the caliper, replacement, which units are matched, decide both what is estimated and which population the estimate describes. Its resemblance to a paired experiment is superficial: the pairs are assembled after treatment occurred, and unconfoundedness is assumed exactly as before.
4 connections
Contrasts with: Blocked and Paired Randomized Experiments
Requires: Propensity Scores, Unconfoundedness and Overlap
Contrasts with: Inverse Probability Weighting
- Neyman Repeated-Sampling Inference
Hold the potential outcomes fixed and let the assignment vary: that is the frame in which a difference in means has a variance at all. The exact design variance contains a term built from both potential outcomes per unit, which no study observes, so the estimator used in practice deliberately drops it. The result is conservative rather than exact, and knowing which it is changes what an interval claims.
8 connections
Contrasts with: Randomized Assignment
Requires: Potential Outcomes and the Fundamental Problem, Randomized Assignment
Suggested next: Blocked and Paired Randomized Experiments
Contrasts with: Fisher Randomization Inference
Required by: Blocked and Paired Randomized Experiments, Regression Adjustment in Experiments
Suggested after: Randomized Assignment
- Potential Outcomes and the Fundamental Problem
Each unit has an outcome under treatment and an outcome under control; the effect for that unit is their difference. Exactly one of the two is ever observed, so no individual effect is ever computed. Every method in causal inference is a way of replacing the missing half with a credible comparison, and every such method is a claim about an average rather than about a person.
7 connections
Suggested next: Randomized Assignment
Contrasts with: Randomized Assignment
Required by: Fisher Randomization Inference, Neyman Repeated-Sampling Inference
- Propensity Scores
The probability of treatment given the covariates is a single number that can stand in for all of them. Conditioning on it balances the covariates it was built from, which turns a high-dimensional adjustment problem into a scalar one. It does nothing whatever about variables nobody measured. Judge a specification by covariate balance and overlap rather than by how well it predicts treatment, which measures something the adjustment does not need.
9 connections
Contrasts with: Randomized Assignment
Requires: Unconfoundedness and Overlap
Suggested next: Inverse Probability Weighting
Contrasts with: Binary Outcome Models for Experimental Research, Regression Adjustment in Experiments
Required by: Inverse Probability Weighting, Matching for Causal Inference
Suggested after: Unconfoundedness and Overlap
Uses: Conditional Probability, Total Probability and Bayes' Rule
- Random Variables, Expectation and Variance
A random variable as a numeric function of an uncertain outcome, its expectation as the probability-weighted balance point of the distribution, and its variance as the weighted average squared deviation from that point, together with the linearity and scaling rules that govern how both behave under a linear transformation.
10 connections
- Randomized Assignment
An assignment mechanism is the chance process that decides which units are treated, and it is a design decision recorded before outcomes are seen. Because it is generated independently of the potential outcomes, it makes treated and control groups comparable in expectation, which is what licenses reading a difference in means as a causal effect. Bernoulli and complete randomization differ in what they hold fixed, and the difference changes the inference that follows.
12 connections
Contrasts with: Potential Outcomes and the Fundamental Problem
Requires: Potential Outcomes and the Fundamental Problem
Suggested next: Neyman Repeated-Sampling Inference
Contrasts with: Blocked and Paired Randomized Experiments, Neyman Repeated-Sampling Inference
and 2 more
Propensity Scores, Unconfoundedness and OverlapRequired by: Blocked and Paired Randomized Experiments, Fisher Randomization Inference
Suggested after: Potential Outcomes and the Fundamental Problem
- Regression Adjustment in Experiments
A regression of outcome on a treatment indicator reproduces the difference in means exactly. Adding pretreatment covariates can sharpen the estimate by absorbing outcome variation the treatment had nothing to do with. What adjustment cannot do is supply identification: randomization already did that, and where randomization is absent no covariate on the right-hand side restores it.
5 connections
- Sampling Distributions and Standard Error
A statistic computed from a sample would have come out differently had the sample been different. The distribution of those hypothetical values is the sampling distribution, and the standard errors, intervals and tests of classical inference are built from it. Its most-used consequence, that the mean of enough observations is approximately normal, is a statement about the mean, not about the data.
7 connections
- Unconfoundedness and Overlap
When nobody assigned treatment, adjustment can still identify an effect, but only under two assumptions. Unconfoundedness says the measured covariates are enough to make assignment as good as random within their levels; overlap says both treatment states are actually possible at every covariate value that matters. They fail differently and, critically, they can be checked differently: overlap is visible in the data, and unconfoundedness is not.
7 connections
Contrasts with: Randomized Assignment
Requires: Potential Outcomes and the Fundamental Problem
Suggested next: Propensity Scores
Required by: Inverse Probability Weighting, Matching for Causal Inference
Financial Mathematics
- Valuing a Stream of Dated Payments
Why a payment's value depends on when it arrives, the closed forms for level and perpetual streams, how competing projects are ranked by net present value at a stated rate, and why the internal rate of return can be non-unique or disagree with that ranking.
5 connections
Analogous to: Optimal Solutions and Optimal Values
Requires: Functions, Domains and Inequalities, Sequences and Their Limits
Used by: The Natural Logarithm
Inferential Statistics
- Estimators and How They Are Judged
An estimator as a random variable rather than a number, the two standard ways of constructing one, and the properties that decide between competitors: bias, variance, their combination as mean squared error, consistency, and why an unbiased estimator is not automatically the better choice.
5 connections
Analogous to: Estimating Out-of-Sample Error
Best taken before: Confidence Intervals for Experimental Research
Requires: Random Variables, Expectation and Variance, Sampling Distributions and Standard Error
Required by: Testing Counts Against a Claim
- Testing Counts Against a Claim
Comparing observed counts with the counts a hypothesis predicts, how the degrees of freedom follow from the table shape and from any parameter estimated along the way, when the chi-square approximation can be trusted, and which procedure remains available when it cannot.
4 connections
Knowledge Representation and the Semantic Web
- Concept Hierarchies as Ordered Structures
Subsumption between concepts as a partial order, the meets and joins that exist among named concepts and the pairs that have none, how a classifier derives subsumptions nobody asserted, and why an ontology's silence differs from a database's negative answer.
4 connections
Contrasts with: Solving and Characterising Linear Systems
Related: Functions, Domains and Inequalities
Requires: Sets and Fields
Related: Mappings That Preserve Structure
Linear Algebra
- Basis and Dimension
A basis is a set that is both spanning and independent, so every vector is a combination of it in exactly one way. Every basis of a given space has the same number of elements, and that number is the dimension, which is what makes dimension a property of the space rather than of a chosen description of it.
9 connections
Related: Linear Independence, Rank, and Bases, Linear Transformations
Requires: Vector Spaces and Subspaces
Related: Sets and Fields
Required by: Determinants, Diagonalization
- Complex Numbers
Numbers of the form
with , forming a field in which every polynomial splits. Arithmetic, conjugation and modulus in rectangular form; rotation and scaling in polar form; and the conjugate-pair structure that governs the eigenvalues of a real matrix.7 connections
Related: Diagonalization, Eigenvalues and Eigenvectors
and 1 more
Vector Spaces and SubspacesRelated: Sets and Fields, Trigonometry
Required by: Second-Order Linear Differential Equations, Systems of Linear Differential Equations
- Cross Products and Geometry in Space
The cross product of two vectors in
is a vector perpendicular to both whose length is the area of the parallelogram they span. With the scalar triple product it measures volume, decides coplanarity, and gives closed-form distances from a point to a line or a plane. 5 connections
Related: The Singular Value Decomposition
Requires: Determinants, Vectors and Linear Combinations
Related: Permutations
Required by: Partial Derivatives, the Gradient and Critical Points
- Determinants
A single number attached to a square matrix, computable by cofactor expansion along any row or column, that vanishes exactly when the columns are dependent. It scales by the same factor a linear map scales volume, and it multiplies across products, which is what makes it useful rather than merely computable.
9 connections
Related: Linear Independence, Rank, and Bases, Linear Transformations
Requires: Basis and Dimension
Related: Diagonalization, Green's Theorem
and 1 more
PermutationsRequired by: Cross Products and Geometry in Space, Eigenvalues and Eigenvectors
and 1 more
Matrix Inverses and Elementary Matrices - Diagonalization
Writing a matrix as
, where the columns of are eigenvectors and holds the eigenvalues. It is possible exactly when a basis of eigenvectors exists, and it turns repeated application of the matrix into arithmetic on numbers. 8 connections
Related: Determinants
Requires: Basis and Dimension, Eigenvalues and Eigenvectors
Contrasts with: The Singular Value Decomposition
Related: Complex Numbers, Sets and Fields
and 1 more
Systems of Linear Differential EquationsRequired by: Quadratic Forms and Definiteness
- Eigenvalues and Eigenvectors
The directions a linear map leaves in place, and the factors by which it scales them. They are found as the roots of
and the null spaces of ; whether enough of them exist to describe the whole space is a separate question with a precise answer.8 connections
Related: Linear Transformations
Requires: Basis and Dimension, Determinants
Related: Complex Numbers
Required by: Diagonalization, Quadratic Forms and Definiteness
- Linear Transformations
A map between vector spaces is linear when it respects addition and scaling. Every such map on finite-dimensional spaces is determined by what it does to a basis, which is what a matrix records; its kernel and image are subspaces whose dimensions sum to the dimension of the domain.
10 connections
Related: Matrices as Operators
Requires: Vector Spaces and Subspaces
Analogous to: Mappings That Preserve Structure
Related: Basis and Dimension, Determinants
- Matrix Inverses and Elementary Matrices
Computing
by Gauss–Jordan elimination, and seeing why it works: each row operation is multiplication by an elementary matrix, so a reduction to the identity is a factorisation of the inverse. The same bookkeeping gives the LU factorisation. 6 connections
Related: Matrices as Operators, The Revised Simplex Method
Requires: Basis and Dimension, Determinants
Related: Permutations, Sets and Fields
- Orthogonality and Projection
An inner product gives a vector space lengths and angles. Orthonormal bases make coordinates into inner products, Gram–Schmidt produces one from any basis, and projecting a vector onto a subspace finds the closest point in it, which is what least squares computes.
6 connections
- Permutations
Rearrangements of
, their cycle structure, and the sign that counts whether a rearrangement is reachable by an even or odd number of swaps. The sign is what the determinant's row-swap rule records and what the Leibniz formula sums over. 3 connections
- Quadratic Forms and Definiteness
A homogeneous degree-two expression written as
with symmetric. The spectral theorem diagonalises it by an orthogonal change of variables, so the eigenvalues of decide whether the form is always positive, always negative, or takes both signs. 6 connections
- Sets and Fields
The set notation linear algebra is written in, and the field axioms its scalars must satisfy. Which familiar systems are fields, which fail and at which axiom, and what every result that says "over a field" is actually assuming.
7 connections
- The Singular Value Decomposition
Every matrix factors as
with orthonormal and and nonnegative diagonal , including rectangular and singular ones, which diagonalization cannot handle. The singular values measure how much the map stretches along each of a set of orthogonal directions, and truncating them gives the best low-rank approximation. 6 connections
Contrasts with: Diagonalization
Requires: Eigenvalues and Eigenvectors, Orthogonality and Projection
Related: Cross Products and Geometry in Space, Quadratic Forms and Definiteness
and 1 more
Shrinkage, and the Trade It Makes - Vector Spaces and Subspaces
A vector space is a set with an addition and a scaling that obey fixed rules. A subspace is a subset that is a vector space in its own right, which three checks decide: it contains the zero vector, and it is closed under addition and under scaling. The checks are separable, and a set can pass one while failing another.
8 connections
Operational Research
- Active Constraints
Standing at a point of a feasible region, some restrictions are pressing and the rest have room to spare. The ones holding with equality form the active set, and they are what pins a point down: how many hold, and whether their normals are independent, decides whether the point is a corner, an edge point, or interior.
8 connections
- Adjacent Basic Solutions
Two bases are adjacent when they differ in exactly one column. Under nondegeneracy the distinct corners they describe are joined by an edge of the feasible region; at a degenerate corner two adjacent bases can name the same point, and the edge between them has length zero. The edge has a direction computable from the basis, and moving along it until a basic variable reaches zero is what carries you from one corner to the next. This is the bridge between the static geometry of corners and the step the simplex method takes.
8 connections
Contrasts with: Degenerate Basic Feasible Solutions
Requires: Active Constraints, Basic Solutions
Suggested next: One Iteration of the Simplex Method
Used by: One Iteration of the Simplex Method
Required by: One Iteration of the Simplex Method
Uses: Active Constraints
- Basic Solutions
Choosing
independent columns of and solving for them, with every other variable set to zero, produces a basic solution. It is a construction, and it can be carried out correctly and still produce a point outside the feasible region: nothing in the procedure tests the sign restrictions. Feasibility is a separate question asked afterwards. 13 connections
Contrasts with: Extreme Points and Basic Feasible Solutions
Requires: Converting a Linear Program to Standard Form, Linear Independence, Rank, and Bases
Suggested next: Extreme Points and Basic Feasible Solutions
Used by: Extreme Points and Basic Feasible Solutions
Required by: Adjacent Basic Solutions, Canonical Form for a Basis
Suggested after: Linear Independence, Rank, and Bases
Uses: The Matrix Form of a Linear Program, Transforming the Variables
- Canonical Form for a Basis
Partitioning a standard-form program by a basis and solving for the basic variables rewrites the whole program in terms of the nonbasic ones. The objective becomes a constant plus a weighted sum of nonbasic variables, and those weights are the reduced costs. The optimality test is then a matter of reading signs off the page rather than computing anything further.
7 connections
- Constrained Optimization
Three parts settle every optimization question: the quantities a decision maker sets directly, a function measuring what counts as better, and the set of choices the limits permit. Linear programming is one instance of this shape, and every method in the subject assumes a problem already put into it.
6 connections
Contrasts with: Formulating a Linear Program
Suggested next: Feasible Sets
Contrasts with: Choices No Utility Function Represents
Required by: Feasible Sets, From a Described Problem to a Model
and 1 more
Optimal Solutions and Optimal Values - Contour Lines and the Direction of Improvement
An objective assigns a value to every point, so points of equal value form a family of parallel lines and the objective's coefficient vector points perpendicular to them, towards increase. Drawing the family correctly and knowing which way along it improves are two different skills, and the second is where the errors are.
9 connections
Contrasts with: The Feasible Region of a Linear Program
Requires: Plotting Linear Inequalities, Vectors and Linear Combinations
Suggested next: Solving a Two-Variable Linear Program Graphically
Used by: Reduced Costs and the Optimality Test, Solving a Two-Variable Linear Program Graphically
Contrasts with: Plotting Linear Inequalities
Required by: Solving a Two-Variable Linear Program Graphically
Suggested after: Plotting Linear Inequalities
- Converting a Linear Program to Standard Form
Standard form requires a minimization objective, equality constraints, and nonnegative variables. Any linear program can be converted to an equivalent standard-form program by negating a maximization objective, introducing slack or surplus variables, and splitting free variables into a difference of two nonnegative variables.
15 connections
Best taken before: Transforming the Variables
Suggested next: Transforming the Variables
Used by: Extreme Points and Basic Feasible Solutions
Contrasts with: The Feasible Region of a Linear Program, Transforming the Variables
Best taken after: Formulating a Linear Program, The Matrix Form of a Linear Program
Required by: Basic Solutions, Canonical Form for a Basis
- Degenerate Basic Feasible Solutions
A basic feasible solution is degenerate when fewer than
of its components are strictly positive. Geometrically, more constraints are active at that corner than are needed to pin it down; algebraically, several different bases describe the same point. The practical consequence is that a basis change can leave the point where it was, which is how the simplex method stalls and, in rare cases, cycles. 12 connections
Contrasts with: Extreme Points and Basic Feasible Solutions
Requires: Active Constraints, Basic Solutions
Suggested next: One Iteration of the Simplex Method
Used by: One Iteration of the Simplex Method
Contrasts with: Adjacent Basic Solutions
Best taken after: Extreme Points and Basic Feasible Solutions
Required by: One Iteration of the Simplex Method
Suggested after: Active Constraints, Extreme Points and Basic Feasible Solutions
Uses: Active Constraints
- Driving a Linear Programming Solver
Getting a formulated program into the argument form a particular solver expects, and reading its output back into the model's own terms. The mathematics is unchanged; what varies is the convention each tool imposes and the places a mismatch produces a confident wrong answer.
3 connections
- Extreme Points and Basic Feasible Solutions
For a linear program in standard form, a point of the feasible region is an extreme point exactly when it is a basic feasible solution. The geometric notion of a corner and the algebraic notion of a basis describe the same set of points, which is why an algorithm that moves between bases is searching corners.
21 connections
Best taken before: Degenerate Basic Feasible Solutions
Requires: Basic Solutions, Converting a Linear Program to Standard Form
Suggested next: Degenerate Basic Feasible Solutions
Used by: Converting a Linear Program to Standard Form
Has parts: Solving and Characterising Linear Systems
Contrasts with: Basic Solutions, Degenerate Basic Feasible Solutions
and 1 more
Integer ProgramsRepresents: Solving a Two-Variable Linear Program Graphically
Required by: Adjacent Basic Solutions, Degenerate Basic Feasible Solutions
Suggested after: Basic Solutions
- Feasible Sets
What survives after every restriction is applied at once, together with two questions that precede any search for an optimum: whether anything survives at all, and whether what survives runs on without limit. Linear programs specialise this abstraction; they do not define it.
7 connections
Generalizes: The Feasible Region of a Linear Program
Requires: Constrained Optimization
Suggested next: Optimal Solutions and Optimal Values
Related: Half-Spaces and Hyperplanes, The Four Terminal Outcomes
Required by: Optimal Solutions and Optimal Values
Suggested after: Constrained Optimization
- Formulating a Linear Program
Formulating a linear program means turning a described decision problem into decision variables, a linear objective, and linear constraints. The hard part is not the algebra but the choices: what exactly is being decided, in what units, and which stated conditions are genuine restrictions rather than commentary.
14 connections
Best taken before: Converting a Linear Program to Standard Form
Suggested next: Solving a Two-Variable Linear Program Graphically
Contrasts with: Constrained Optimization, Modelling with Binary Variables
and 1 more
Recurring Shapes of Linear ProgramsBest taken after: From a Described Problem to a Model
Required by: Driving a Linear Programming Solver, Integer Programs
Uses: From a Described Problem to a Model, Verifying a Reported Solution
- From a Described Problem to a Model
A description in words becomes a model by answering three questions in order: what is chosen, what is being optimised, and what limits the choice. Each answer has a syntactic form, and the order matters because the objective and the constraints are both written in terms of the variables. Most modelling faults are traceable to a variable that was never pinned down.
3 connections
Best taken before: Formulating a Linear Program
Requires: Constrained Optimization
Used by: Formulating a Linear Program
- Half-Spaces and Hyperplanes
A single linear equation describes a flat boundary that divides space in two, and a single linear inequality describes one of the two sides together with that boundary. These are the pieces every feasible region is assembled from: the region is what survives when all the allowed sides are intersected.
8 connections
Part of: The Feasible Region of a Linear Program
Related: Feasible Sets
Suggested next: The Feasible Region of a Linear Program
Has parts: Vectors and Linear Combinations
Related: The Row and Column Pictures
Required by: Active Constraints, Plotting Linear Inequalities
and 1 more
The Feasible Region of a Linear Program - Integer Programs
An integer program is a linear program with some variables restricted to whole numbers. The restriction looks small and changes everything: the feasible set becomes a scatter of isolated points rather than a polyhedron, its optimum need not sit at a corner, and the geometric account that makes the simplex method work no longer applies.
10 connections
Contrasts with: Extreme Points and Basic Feasible Solutions
Requires: Formulating a Linear Program, The Feasible Region of a Linear Program
Suggested next: Modelling with Binary Variables
Used by: Modelling with Binary Variables, The Linear Relaxation and Its Bound
Contrasts with: The Linear Relaxation and Its Bound, The Transportation Model
Required by: Modelling with Binary Variables, The Linear Relaxation and Its Bound
- Linear Independence, Rank, and Bases
A set of vectors is linearly independent when none of them is a combination of the others. The rank of a matrix is the number of independent columns it has, and a square matrix is invertible exactly when its columns are independent. These facts decide which column sets can serve as a basis, and therefore which points a linear program can call a basic solution.
11 connections
Suggested next: Basic Solutions
Used by: Extreme Points and Basic Feasible Solutions, One Iteration of the Simplex Method
Has parts: Matrices as Operators, Vectors and Linear Combinations
Related: Basis and Dimension, Determinants
and 1 more
Vector Spaces and SubspacesRequired by: Basic Solutions, Extreme Points and Basic Feasible Solutions
and 1 more
One Iteration of the Simplex Method - Matrices as Operators
A matrix is best understood by what it does: applied to a vector it returns a linear combination of its own columns, weighted by that vector's entries. Reading
that way turns a wall of arithmetic into one idea, and it is the reading every later result depends on. A constraint system, a basis, a pivot step. Multiplication of matrices is then composition of those actions, which is why order matters and why it is not commutative. 8 connections
Part of: Linear Independence, Rank, and Bases
Requires: Vectors and Linear Combinations
Suggested next: Solving and Characterising Linear Systems
Related: Linear Transformations, Matrix Inverses and Elementary Matrices
Required by: Solving and Characterising Linear Systems, The Matrix Form of a Linear Program
Suggested after: Vectors and Linear Combinations
- Modelling with Binary Variables
A variable restricted to zero or one records a yes-or-no decision, and a small vocabulary of linear constraints turns logical conditions into algebra: at most one of these, if this then that, this only if that is open. The patterns are few and compose, and the recurring error is writing a condition that is true of the situation but does not constrain the model.
6 connections
Contrasts with: Formulating a Linear Program
Requires: Formulating a Linear Program, Integer Programs
Used by: The Linear Relaxation and Its Bound
Suggested after: Integer Programs
Uses: Integer Programs
- One Iteration of the Simplex Method
A simplex iteration moves from one basic feasible solution to an adjacent one. A nonbasic variable with negative reduced cost enters the basis; the minimum ratio test decides how far it can increase before a basic variable reaches zero, and that variable leaves. The test is what keeps the new point feasible.
16 connections
Requires: Adjacent Basic Solutions, Degenerate Basic Feasible Solutions
Used by: Solving a Two-Variable Linear Program Graphically
Contrasts with: Verifying a Reported Solution
Required by: The Full Tableau Simplex Method
Suggested after: Adjacent Basic Solutions, Degenerate Basic Feasible Solutions
and 1 more
Reduced Costs and the Optimality Test - Optimal Solutions and Optimal Values
An optimization problem yields two different objects: the choice that achieves the best outcome, and the number that outcome is worth. Several choices may tie for best, so the first can be plural while the second never is, and a problem can have feasible points yet attain neither.
4 connections
Requires: Constrained Optimization, Feasible Sets
Analogous to: Valuing a Stream of Dated Payments
Suggested after: Feasible Sets
- Plotting Linear Inequalities
A linear inequality is drawn in two moves: draw its boundary line, then decide which of the two sides it allows. The first is arithmetic and rarely goes wrong. The second has a failure mode of its own, the wrong side chosen, which survives a perfectly accurate drawing and quietly produces a region that is not the feasible one.
7 connections
Contrasts with: Contour Lines and the Direction of Improvement
Requires: Half-Spaces and Hyperplanes
Suggested next: Contour Lines and the Direction of Improvement
Used by: Solving a Two-Variable Linear Program Graphically, The Feasible Region of a Linear Program
Required by: Contour Lines and the Direction of Improvement, Solving a Two-Variable Linear Program Graphically
- Recurring Shapes of Linear Programs
Most described decision problems fall into a few shapes: meet requirements at least cost, use limited resources for the most contribution, or divide a fixed budget under exposure limits. Recognising the shape fixes the direction of the objective, the direction of the constraints, and which terminal outcome is the one to watch for, before any coefficient is written.
4 connections
Contrasts with: Formulating a Linear Program
Related: The Four Terminal Outcomes
Requires: Formulating a Linear Program
Required by: The Transportation Model
- Reduced Costs and the Optimality Test
The reduced cost of a nonbasic variable is the net change in the objective per unit increase of that variable, once the basic variables adjust to keep the constraints satisfied. For a minimization program in standard form, a basic feasible solution is optimal when every reduced cost is nonnegative, because no available direction improves the objective.
7 connections
Requires: Converting a Linear Program to Standard Form, Extreme Points and Basic Feasible Solutions
Suggested next: One Iteration of the Simplex Method
Required by: One Iteration of the Simplex Method, The Revised Simplex Method
Uses: Canonical Form for a Basis, Contour Lines and the Direction of Improvement
- Solving a Two-Variable Linear Program Graphically
A linear program in two variables can be solved by drawing: shade the feasible region, draw the objective as a family of parallel contour lines, and push the contour in the improving direction until it is about to leave the region. The last point it touches is optimal, and it is always a vertex unless a whole edge is touched at once.
14 connections
Best taken before: The Four Terminal Outcomes
Represented by: Extreme Points and Basic Feasible Solutions
Requires: Contour Lines and the Direction of Improvement, Formulating a Linear Program
Suggested next: The Four Terminal Outcomes
Required by: The Four Terminal Outcomes
Suggested after: Contour Lines and the Direction of Improvement, Formulating a Linear Program
and 1 more
The Feasible Region of a Linear ProgramUses: Contour Lines and the Direction of Improvement, One Iteration of the Simplex Method
and 1 more
Plotting Linear Inequalities - Solving and Characterising Linear Systems
A linear system
has exactly one of three outcomes: no solution, exactly one, or infinitely many. Elimination finds them, and the shape of the reduced system says which case holds and why. Two of those cases occur in linear programming: an infeasible program is a system with no solution, and a program with many optima has a solution set with a free direction.7 connections
Part of: Extreme Points and Basic Feasible Solutions
Requires: Matrices as Operators
Suggested next: The Row and Column Pictures
Contrasts with: Concept Hierarchies as Ordered Structures
Required by: Basic Solutions, The Row and Column Pictures
Suggested after: Matrices as Operators
- The Feasible Region of a Linear Program
The feasible region of a linear program is the set of points satisfying every constraint simultaneously. It is the intersection of finitely many half-spaces, which makes it a convex polyhedron: possibly empty, possibly unbounded, and never containing a dent.
15 connections
Contrasts with: Converting a Linear Program to Standard Form
Requires: Half-Spaces and Hyperplanes
Suggested next: Solving a Two-Variable Linear Program Graphically
Has parts: Half-Spaces and Hyperplanes, The Row and Column Pictures
Contrasts with: Contour Lines and the Direction of Improvement, The Four Terminal Outcomes
Generalized by: Feasible Sets
Required by: Active Constraints, Extreme Points and Basic Feasible Solutions
Suggested after: Half-Spaces and Hyperplanes
- The Four Terminal Outcomes
Every linear program ends in exactly one of four states: infeasible, unbounded, a unique optimum, or infinitely many optima. Deciding which one holds is a single competence, and it cannot be practised one case at a time. The work is telling them apart, and each has a characteristic signature in the picture, in the algebra, and in what a solver reports.
12 connections
Contrasts with: The Feasible Region of a Linear Program
Related: Feasible Sets
Requires: Solving a Two-Variable Linear Program Graphically, The Feasible Region of a Linear Program
Suggested next: Verifying a Reported Solution
Used by: One Iteration of the Simplex Method
Best taken after: Solving a Two-Variable Linear Program Graphically
Related: Recurring Shapes of Linear Programs, The Full Tableau Simplex Method
Required by: The Linear Relaxation and Its Bound, Verifying a Reported Solution
Suggested after: Solving a Two-Variable Linear Program Graphically
- The Full Tableau Simplex Method
Running the simplex method to termination in a single table. The tableau carries the whole system in canonical form against the current basis, each pivot updates it by row operations, and the stopping conditions, optimality and unboundedness, are read off the table rather than computed separately.
4 connections
Related: The Four Terminal Outcomes
Requires: Canonical Form for a Basis, One Iteration of the Simplex Method
Contrasts with: The Revised Simplex Method
- The Linear Relaxation and Its Bound
Drop the integrality restriction and an integer program becomes a linear program that can actually be solved. Its optimal value is a bound on the integer optimum, never worse, in a direction fixed by whether you are minimising or maximising, and that bound is the foundation of every serious integer method. What the relaxation does not give you is the answer: rounding its solution is not generally valid, and often not even feasible.
6 connections
Contrasts with: Integer Programs
Requires: Integer Programs, The Four Terminal Outcomes
Used by: Verifying a Reported Solution
- The Matrix Form of a Linear Program
Collecting the coefficients of a linear program into a matrix and its data into vectors replaces
written constraints with one statement, . The compression is not cosmetic: the row reading recovers the individual constraints, the column reading exhibitsas a combination of the columns of , and every later method is stated in terms of one reading or the other. 7 connections
Analogous to: The Row and Column Pictures
Best taken before: Converting a Linear Program to Standard Form
Requires: Formulating a Linear Program, Matrices as Operators
Used by: Basic Solutions
Required by: Canonical Form for a Basis, The Transportation Model
- The Revised Simplex Method
The same algorithm organised around the basis inverse rather than a full table. Multipliers are formed once per iteration, columns are priced only as they are examined, and the single direction needed for the ratio test is computed on demand. What changes is what gets stored and recomputed, not which basis sequence the method visits.
4 connections
Contrasts with: The Full Tableau Simplex Method
Related: Canonical Form for a Basis
Requires: Reduced Costs and the Optimality Test
- The Row and Column Pictures
One system, two pictures. Read by rows,
asks where several flats intersect. Read by columns, it asks whethercan be mixed from the columns of . Both describe the same solutions, and each makes visible what the other hides, which is why moving between them is a competence rather than a preference. 5 connections
Part of: The Feasible Region of a Linear Program
Related: Half-Spaces and Hyperplanes
Requires: Solving and Characterising Linear Systems
Analogous to: The Matrix Form of a Linear Program
Suggested after: Solving and Characterising Linear Systems
- The Transportation Model
Shipping a single commodity from supply points to demand points at least cost. The model has one structural requirement, total supply must equal total demand, and two readings, a cost table and a bipartite network, which answer different questions about the same program.
3 connections
Contrasts with: Integer Programs
Requires: Recurring Shapes of Linear Programs, The Matrix Form of a Linear Program
- Transforming the Variables
Slack and surplus variables change the form of a constraint. A different family of substitutions changes the variables themselves: splitting a sign-unrestricted variable into a difference of two nonnegatives, shifting a variable with a nonzero lower bound, and negating a nonpositive one. Each preserves the feasible set and each carries its own error, the most consequential being that the split is not unique.
5 connections
Contrasts with: Converting a Linear Program to Standard Form
Requires: Converting a Linear Program to Standard Form
Used by: Basic Solutions
Best taken after: Converting a Linear Program to Standard Form
Suggested after: Converting a Linear Program to Standard Form
- Vectors and Linear Combinations
A vector is a list of numbers that behaves like a direction with a length. Adding two of them and scaling one are the only operations linear algebra is built from, and every later object, a constraint, a cost, a basis, is assembled out of those two moves. The dot product turns that geometry into arithmetic: it measures how much of one direction lies along another, which is what lets an optimiser decide where to move without drawing anything.
8 connections
Part of: Half-Spaces and Hyperplanes, Linear Independence, Rank, and Bases
Suggested next: Matrices as Operators
Related: Orthogonality and Projection, Vector Spaces and Subspaces
Required by: Contour Lines and the Direction of Improvement, Cross Products and Geometry in Space
and 1 more
Matrices as Operators - Verifying a Reported Solution
A solver returns a status, a point and a value. Checking that answer is a competence separate from producing one, and it is cheap: substitute the point into every constraint, recompute the objective, and ask whether the reported status is consistent with what you find. Each check catches a different class of fault, and a model that solves without error can still be answering a different question than the one you posed.
7 connections
Contrasts with: One Iteration of the Simplex Method
Requires: Converting a Linear Program to Standard Form, The Four Terminal Outcomes
Used by: Formulating a Linear Program
Required by: Driving a Linear Programming Solver
Suggested after: The Four Terminal Outcomes
Psychometrics
- Internal Consistency and Coefficient Alpha
Computing coefficient alpha from item and total-score variances, the model assumptions that decide whether it equals reliability, understates it, or can exceed it, and why the coefficient rises with test length whatever the items measure and reports nothing about how many dimensions they span.
2 connections
- Reliability and Measurement Error
An observed score as a true score plus error, reliability as the share of observed variance that is not error, the standard error of measurement that puts the same information on the score scale, and why both are properties of scores from a population rather than of the instrument.
- Validity as an Argument for an Interpretation
Validity as the degree to which evidence supports a specific interpretation of scores for a specific use, the kinds of evidence that bear on such a claim, why a consistent score can be consistently measuring the wrong thing, and why evidence gathered for one use does not transfer to another.
3 connections
Requires: Reliability and Measurement Error
Analogous to: Choices No Utility Function Represents
Statistical Learning
- Association Rules, and What Confidence Leaves Out
How support, confidence and lift are computed from transaction counts, why the anti-monotone property lets a level-wise search examine a fraction of the candidate lattice, and the case that matters most: a rule whose confidence passes any conventional threshold while its lift falls below one, so the consequent is less frequent among transactions containing the antecedent than it is overall, which is the opposite of what the confidence figure suggests.
2 connections
- Choosing a Regression Form
What the response variable and the shape of the relationship each rule out, how polynomial terms curve a fit while leaving it linear in its coefficients, why exponential and Poisson models are read on a multiplicative scale, and what a fitted form does beyond the range it was fitted on.
7 connections
Requires: Estimating Out-of-Sample Error, Linear Regression for Experimental Research
Used by: The Natural Logarithm
Contrasts with: Fitting Without a Formula, Trees and Ensembles of Them
Related: Shrinkage, and the Trade It Makes
- Clustering, and What the Objective Assumes
How k-means alternates assignment and update until the labels settle, why a converged solution can be far from the best one and how restarts detect that, what a within-cluster sum of squares curve does and does not tell you about the number of groups, and the case that matters most: concentric rings, where the ring grouping scores worse on the objective than the partition the method returns, so the failure lies in what the objective treats as a cluster rather than in the algorithm.
5 connections
Contrasts with: Estimating Out-of-Sample Error
Related: Trees and Ensembles of Them
Requires: Random Variables, Expectation and Variance
Contrasts with: Linear Discriminants, and the Covariance They Assume
- Estimating Out-of-Sample Error
Why the error a model reports on the data it was fitted to is not the error it will make on new data, how resampling estimates the second, what bias, variance and irreducible noise each contribute, and which splitting schemes remain valid when the observations are ordered in time or reused for tuning.
13 connections
Related: Sampling Distributions and Standard Error
Requires: Covariance, Independence and the Variance of a Sum, Linear Regression for Experimental Research
Analogous to: Estimators and How They Are Judged
Contrasts with: Clustering, and What the Objective Assumes, When a Coefficient Is Not an Effect
Related: Choosing What the Reader Will Judge
Required by: Choosing a Regression Form, Fitting Without a Formula
- Fitting Without a Formula
Estimating a response at a point by averaging nearby observations rather than by fitting a formula, how the neighbourhood size or bandwidth trades bias against variance, why training error cannot choose it, and what disappears when no functional form is assumed.
3 connections
Contrasts with: Choosing a Regression Form
Related: Trees and Ensembles of Them
Requires: Estimating Out-of-Sample Error
- Linear Discriminants, and the Covariance They Assume
Why the best direction for separating two labelled classes is generally not the line joining their means, how dividing by the pooled within-class scatter rotates it toward directions in which the classes are internally tight, and what happens when the single-covariance assumption behind that pooling fails. A case where the method returns a boundary, reports nothing amiss, and classifies barely better than chance.
5 connections
- Shrinkage, and the Trade It Makes
Why least squares returns wild coefficients when two predictors carry nearly the same information, how adding
to the diagonal of steadies them, what the lasso does differently in setting coefficients to exactly zero, and why a penalty is a trade rather than a repair. Least squares already minimises the unpenalised training sum of squares, so no penalised fit can have a lower training sum of squares, and in practice it is higher. The case for a penalty is therefore made on held-out error. - Trees and Ensembles of Them
How a regression tree chooses its splits by exhaustive search, why the resulting fit is piecewise constant and unstable under resampling, and what averaging many trees changes about that instability, including the correlation floor that limits what more trees can achieve and the restriction random forests use to reduce it.
5 connections