In a standard feed-forward network, composing affine layers without a nonlinear operation collapses to another affine transformation. Nonlinear activations let the network represent functions that a single affine layer cannot. Activations also shape gradient flow, output constraints, sparsity, numerical behavior, and compatibility with initialization and normalization.