✦ For everyone, free.

Practical knowledge for real and everyday life

Home

Complexity and Regularity Descriptors

Complexity and Regularity Descriptors measure signal irregularity and structure, revealing insights into dynamic behaviors and underlying patterns.

Complexity and Regularity Descriptors are explicit characterizations of sequential regularity, pattern persistence, ordinal or symbolic organization, compressibility, recurrence of states, multiscale structure, scaling behavior, fractal-like organization, local divergence, and intrinsic predictability in declared Behavioral Signal evidence. These descriptors are mathematical tools designed to quantify specific features of temporal or sequential data reflecting behavior, yet the terms regularity, irregularity, entropy, randomness, complexity, nonlinearity, fractal scaling, predictability, and chaos are not synonyms. Different families of descriptors operationalize different mathematical properties and can rank the same signals differently. A larger value in one descriptor does not universally mean that the signal is more complex, less regular, healthier, more adaptive, more random, or more behaviorally sophisticated. Each descriptor must be interpreted in the context of its formal definition, assumptions, parameters, and intended operational meaning.


Meaning and Conceptual Boundaries

A Complexity or Regularity Descriptor is a reproducible characterization whose meaning depends on ordered signal structure, pattern construction, state recurrence, scaling across resolutions, symbolic parsing, reconstructed-state geometry, divergence, or prediction behavior under declared parameters and assumptions. Many such descriptors are useful finite-data statistics that do not imply that the underlying behavioral process is nonlinear, deterministic, low-dimensional, fractal, or chaotic. Rather, they quantify operational aspects of the signal or its transformations as defined by the user’s analytical choices.

Regularity concerns recurrence, repeatability, predictability, or restricted pattern variation under a declared definition. Complexity can refer to richer organization across patterns, scales, states, or predictive structure and has no single universally accepted scalar definition. For example, highly periodic signals can be very regular but structurally simple, while white-noise-like signals can be highly irregular yet lack the organized multiscale structure intended by some complexity concepts.

Randomness and stochasticity concern generative uncertainty or inherent unpredictability in signal generation. Nonlinearity indicates the failure of an adopted linear relationship or null model. Deterministic irregularity can arise without stochastic randomness, producing signals that appear irregular but are generated by deterministic rules. Chaos requires a deterministic dynamical interpretation with sensitive dependence on initial conditions and additional evidential conditions. Entropy-like, recurrence, fractal, or divergence descriptors can inform these questions but do not establish them by name alone.

ConceptWhat It MeansRepresentative EvidenceWhat Cannot Be Inferred Automatically
RegularityRepeatability or restricted variation of patterns under a declared ruleLow Approximate Entropy, high determinism in recurrenceThat the signal is simple or non-complex
IrregularityDeparture from repeatability or persistenceHigh Approximate Entropy, low recurrence rateThat the signal is complex or meaningful
Entropy-Like UnpredictabilityUncertainty in pattern occurrence or symbol distributionPermutation Entropy, Sample EntropyThat the underlying process is stochastic or nonlinear
Randomness/StochasticityGenerative uncertainty or noiseStatistical tests, stochastic modelsThat the signal is unpredictable in finite samples
ComplexityRich structure across multiple patterns, scales, or statesMultiscale entropy profiles, fractal scalingThat complexity is scalar or monotonic with descriptor values
NonlinearityFailure of linear model assumptionsNonlinear prediction error, surrogate data rejectionThat the process is chaotic or deterministic
Fractal/Scaling StructureSelf-similar or scale-invariant organizationDFA scaling exponent, fractal dimensionThat the signal is fractal in a strict geometric or dynamical sense
ChaosDeterministic irregularity with sensitive dependence on initial conditionsPositive Lyapunov exponent estimates, recurrence patternsThat chaos is proven by single descriptor or finite-data estimate

Order sensitivity and domain boundaries are critical. A histogram entropy describes the value distribution without temporal order; spectral entropy describes the spectral energy distribution; time-frequency entropy describes localized spectral-temporal representations; Sample Entropy and Permutation Entropy use sequence-pattern structure; event recurrence describes repeated event occurrence times; state-space recurrence describes returns to neighborhoods in a reconstructed or observed state space. The shared words entropy or recurrence do not make these quantities interchangeable.


Template-Matching Regularity and Entropy-Like Descriptors

A generic delayed template vector is constructed as

v i m , τ = [ xi , xi+τ , , xi+(m1)τ ]

where x_i is observation i of the declared ordered scalar series, i is the template-start index, m is the positive integer template dimension, τ is the positive integer sample delay or another explicitly mapped delay, and v_i^(m,τ) is the resulting template vector. The same mathematical vector construction can serve pattern matching or state-space reconstruction but does not automatically represent the complete physical or behavioral state of a person.

Approximate Entropy is a template-matching regularity statistic based on logarithmic match frequencies under declared template dimension, delay, distance metric, tolerance, self-match convention, normalization, and finite-sample rules. It includes self matches in its classical construction and can depend materially on record length and parameter choices, often producing biased estimates if parameters or data length are not carefully controlled.

A representative Sample Entropy abstraction is

SampEn (m,r) = ln ( A B )

where m is the template dimension, r is the declared matching tolerance under a declared distance and scaling convention, B is the eligible count or consistently normalized frequency of distinct template pairs that match at dimension m, and A is the corresponding eligible count or frequency of those pair relationships that remain matched when extended to dimension m+1. Exact counting, delay, edge eligibility, self-match exclusion, and normalization conventions must be documented. Values of A = 0, B = 0, or too few matches can produce infinite, undefined, or unstable results rather than a valid zero.

Sample Entropy is a self-match-excluding template-continuation statistic related to regularity, not an exact universal entropy rate. It is sensitive to parameter choices including m, r, delay, distance metric, support length, amplitude scaling, noise, missingness, and preprocessing. Larger m or smaller r can drastically reduce match counts, and canonical-looking parameter values should not be transferred across modalities without justification.

Fuzzy or soft-matching entropy variants conceptually replace hard binary matches with graded similarity, reducing discontinuity with respect to small distance changes. However, membership function choice, exponent, local/global mean removal, tolerance scaling, and normalization are integral to descriptor identity. Fuzzy Entropy is not one universally standardized formula and requires explicit definition.

DescriptorPattern RuleSelf-Match TreatmentPrimary ParametersMajor Finite-Data Risk
Approximate EntropyBinary match within toleranceIncludes self matchesm, r, τ, tolerance, distanceBias from self-match inclusion; parameter sensitivity
Sample EntropyBinary match within toleranceExcludes self matchesm, r, τ, tolerance, distanceZero matches produce infinite/undefined values
Fuzzy/Soft-Matching EntropyGraded similarity membershipVaries by membership functionMembership parameters, m, r, τParameter dependence, unclear standardization
Matching-Based VariantVariant matching criteriaDepends on variantVaries by variantInterpretation depends on variant specifics

Ordinal, Symbolic, and Compression-Based Descriptors

Normalized Permutation Entropy is defined as

HPE = π Π p(π)ln(p(π)) ln(M!)

where M ≥ 2 is the ordinal-pattern order, Π is the set of admissible ordinal patterns under the declared construction and tie policy, π is one ordinal pattern, p(π) is its empirical relative frequency among eligible delayed windows, and H_PE is Shannon permutation entropy normalized by ln(M!) under the conventional full-pattern normalization. Delay, ties, missing patterns, finite support, and alternative normalizations can change the result.

Permutation Entropy measures the entropy of ordinal-pattern frequencies and thus characterizes local ordering structure rather than amplitude distribution. Embedding order, delay, normalization, overlapping patterns, finite-data coverage of the M! pattern space, and invariance to strictly monotonic amplitude transformations under standard ordinal construction are important considerations. High Permutation Entropy does not prove stochastic randomness or chaos.

Ordinal ties and quantization arise from true plateaus, sensor quantization, rounding, clipping, or repeated states. Tie policies such as order of occurrence, equal-rank patterns, exclusion, deterministic perturbation, or another declared rule produce different ordinal distributions. Random jitter should not be inserted invisibly when exact reproducibility matters.

Symbolic-sequence descriptors arise after explicit symbolization of continuous or discrete evidence. Symbolization can use thresholds, quantiles, ordinal patterns, domain-defined states, or adaptive rules and support word frequencies, block entropy, transition diversity, grammar-like restrictions, and parsing statistics. Symbolization is lossy; its alphabet, thresholds, mapping, missing-symbol rule, and temporal spacing are part of descriptor identity.

Lempel–Ziv-like and compression-related descriptors characterize novelty, phrase growth, or compressibility in a declared symbolic sequence under a specified parsing algorithm and normalization. Raw phrase count, length-normalized variants, alphabet-normalized variants, and actual compressor ratios are not interchangeable. Finite sequence length or symbolization can dominate apparent complexity.

DescriptorInput ConstructionProperty CharacterizedKey ParameterPrimary Misinterpretation
Permutation EntropyOrdinal patterns from scalar seriesLocal ordering complexityEmbedding order M, delay τHigh values imply randomness or chaos
Symbolic/Block EntropySymbolized sequenceSymbol pattern diversityAlphabet size, block lengthThat symbolization preserves all signal info
Lempel–Ziv-Like ParsingSymbolized sequenceNovelty and compressibilityParsing algorithm, normalizationCompression ratio as universal complexity
Compression RatioSymbolized or raw sequenceOverall compressibilityCompressor type, sequence lengthDirect equivalence to complexity
Ordinal-Pattern DiversityOrdinal patternsVariety of ordinal patternsPattern order, tie ruleConfusion with entropy or randomness

Multiscale Complexity and Scale Profiles

Multiscale complexity analysis evaluates a regularity or complexity statistic across a family of explicitly constructed temporal scales rather than treating one single-scale value as the complete system complexity. Conventional coarse-graining, filtering/downsampling, composite or refined-composite constructions, scale-dependent ordinal methods, and other variants exist. The scale-construction rule materially defines the descriptor output and interpretation.

Increasing scale changes temporal grain, effective sampling interval, frequency content, variance, series length, match counts, ordinal-pattern counts, and estimator stability. Higher scales usually contain fewer effective observations, so an apparently smooth complexity profile can become invalid or unstable where support is insufficient.

Retention of the full complexity-versus-scale profile is recommended, with optional summaries such as mean over a declared scale range, area or sum across valid scales, slope, crossover scale, selected-scale vector, or another explicit functional. A scalar complexity index discards scale-specific structure and should not replace the profile when differences across scales are scientifically important.

Multiscale interpretation requires caution. Uncorrelated random noise can be highly irregular at one scale yet lose structure rapidly with coarse-graining, while correlated or structured processes can retain nontrivial organization across scales. Do not interpret monotonic increase or decrease of an entropy profile as universally better or more complex without a declared theoretical meaning.

MethodScale ConstructionOutputData-Length EffectCritical Interpretation Risk
Single-Scale StatisticNone or fixed scaleScalar valueStable with sufficient dataOverlooks scale-dependent structure
Conventional Multiscale EntropyCoarse-graining by averagingProfile of entropy across scaleReduced data length at larger scalesMisinterpretation of scale effects
Composite/Refined-Composite VariantMultiple coarse-grained series combinedSmoothed, more stable profileMore robust to short dataComplexity from composite process, not original
Multiscale Permutation EntropyMultiscale ordinal patternsScale-dependent ordering entropySensitive to pattern coverageOverinterpretation of monotonic trends
Scale-Profile SummaryFunctional reduction of profileScalar summary (mean, slope)Loss of scale-specific infoLoss of critical multiscale information

State-Space Reconstruction and Recurrence Descriptors

A generic recurrence matrix is defined as

R i,j = I ( d ( zi , zj ) ε )

where i and j are eligible state-vector indices, z_i and z_j are observed or reconstructed state vectors under a declared state construction, d is the declared state-space distance or metric, ε is the declared recurrence threshold or radius, I is an indicator function, and R_{i,j} is binary recurrence status for that eligible pair. Temporally adjacent pairs can require exclusion through a declared Theiler or temporal-exclusion rule. Fixed threshold and fixed recurrence-rate constructions define different recurrence analyses.

Delay-coordinate state-space reconstruction from a scalar series is a mathematical construction distinct from direct observation of a complete physical state. Embedding dimension, delay, coordinate scaling, metric, temporal exclusion, missingness, and preprocessing determine neighborhood geometry. A selected embedding dimension should not be interpreted directly as the number of physiological or behavioral degrees of freedom.

Recurrence plots and recurrence quantification characterize returns of observed or reconstructed states to declared neighborhoods. Recurrence Rate quantifies density of recurrence points; Determinism quantifies diagonal-line structure; Average/Max Diagonal Length measure time of predictable evolution; Laminarity quantifies vertical-line structure; Trapping Time estimates average time trapped in states; Recurrence-Time Descriptors summarize intervals between recurrences; Recurrence-Line Entropy quantifies diversity of line lengths. These descriptors quantify different recurrence structures and should not be collapsed into one generic recurrence complexity value.

Threshold, metric, embedding, normalization, and recurrence-density dependence are critical. A fixed distance threshold compares recurrence under one geometric radius, whereas choosing thresholds to force equal recurrence rates removes recurrence density as a between-case difference and changes the interpretation of line-based RQA measures. The construction policy must remain explicit.

State-space recurrence differs from repeated event occurrence. An event-return interval is based on recurrence of declared events in time; an RQA recurrence point is based on closeness of state vectors under a metric and threshold. Neither form should substitute for the other because they answer different scientific questions.

DescriptorRecurrence StructureDefining ParameterInterpretive Caution
Recurrence RateDensity of recurrence pointsRecurrence threshold εDepends on threshold choice; not equivalent to complexity
DeterminismDiagonal-line proportionMinimum line lengthAffected by noise and embedding; not sole predictability measure
Diagonal-Line LengthDuration of predictable segmentsLine length thresholdInfluenced by data length and noise
LaminarityVertical-line proportionMinimum vertical line lengthInterpretation varies with signal type
Trapping TimeAverage vertical line lengthSame as laminaritySensitive to embedding and threshold
Recurrence-Time DescriptorDistribution of recurrence intervalsTemporal exclusionDifferent from event return intervals
Recurrence-Line EntropyDiversity of line lengthsLine length frequency binsDoes not imply stochastic complexity

Scaling, Fractal, and Long-Range Structure

The core Detrended Fluctuation Analysis (DFA) scaling relation is

F ( s ) C s α

where s is the detrending window or scale size under the declared DFA construction, F(s) is the resulting root-mean-square detrended fluctuation at scale s, C is a positive fitted scale factor, and α is the fitted scaling exponent over an explicitly declared log–log scaling range. Detrending order, scale range, crossover structure, weighting, fit diagnostics, and finite record length are part of the estimate. α should not automatically be relabeled a Hurst exponent outside assumptions that justify that relationship.

Hurst-like, rescaled-range, DFA, and related long-range scaling descriptors are distinct estimator families. Persistent, uncorrelated-like, and antipersistent interpretations depend on the process class and estimator assumptions. Trends, periodicities, crossovers, nonstationarity, filtering, and finite support can distort fitted exponents. The actual estimator name and fit range must be preserved rather than reporting a generic fractal exponent.

Fractal-dimension descriptors conceptually differ. Graph-based estimators such as Higuchi-like fractal dimension characterize the signal graph's geometric complexity, distinct from state-space correlation dimension or geometric fractal dimension of a spatial object. Graph-based estimators depend on scale range and sample geometry; correlation dimension depends on reconstructed state space, neighborhood scales, embedding, and evidence of a scaling region. Similar numerical dimensions from different constructions do not imply the same dynamical object.

Multifractal descriptors require sufficient data length and scaling evidence. They distinguish a single scaling exponent from a spectrum of local scaling exponents or generalized dimensions and require scale range, moment orders, regression or partition construction, and uncertainty quantification. A wide estimated multifractal spectrum can be produced or distorted by finite samples, trends, heavy-tailed distributions, noise, or estimator bias and is not automatic evidence of rich behavioral organization.

DescriptorMathematical ObjectRequired ConstructionMain AssumptionMajor Failure Mode
DFA ExponentFluctuation scaling exponent αDetrended fluctuation analysisStationarity within scale rangeCrossover, trends, insufficient data
Hurst-Like ExponentRescaled range or related measureRescaled range or wavelet-basedSelf-similarity, fractional noiseNonstationarity, finite sample bias
Higuchi-Like Fractal DimensionGraph fractal dimensionGraph-based estimatorScale-invariant graph structureScale range misselection, noise
Correlation DimensionState-space attractor dimensionState-space reconstruction, correlation sumsExistence of attractor, scaling regionNo scaling region, embedding errors
Multifractal Spectrum SummarySpectrum of scaling exponentsMultifractal partition functionMultifractal processFinite data, trends, noise distortion

Divergence, Predictability, and Evidence for Nonlinear Structure

Lyapunov-like and local-divergence descriptors estimate how nearby reconstructed trajectories separate over time under declared state-space reconstruction, neighbor-selection rule, temporal exclusion, divergence statistic, time units, and fitted region. Variants include global largest Lyapunov exponent estimates, finite-time or local divergence, mean divergence curves, and inverse-longest-recurrence-line proxies. A positive finite-data estimate is not sufficient proof of deterministic chaos.

Intrinsic or local predictability descriptors such as nearest-neighbor forecast error, horizon-dependent prediction error, local analog forecasting, or improvement over a declared baseline model characterize predictability under a specific prediction rule and state construction. These descriptors differ from supervised prediction of behavioral labels and do not establish causality.

Surrogate-data-supported nonlinear evidence tests whether a chosen complexity or dynamical statistic is inconsistent with a declared null process while preserving selected properties. The null hypothesis, surrogate-generation method, preserved statistics, number of surrogates, randomization seed (when relevant), discriminating statistic, and decision rule must be specified. Rejection of a linear stochastic null is evidence against that null, not proof of deterministic chaos or identification of the true generating mechanism.

Entropy-rate, conditional-entropy, excess-entropy, statistical-complexity, and predictive-information-like quantities have theoretical or estimator-specific meanings different from finite-template SampEn, Permutation Entropy, recurrence-line entropy, spectral entropy, or histogram entropy. Do not rename a convenient finite-data statistic as one of these information-theoretic quantities without the estimator and assumptions required by that definition.


Adequacy, Uncertainty, Sensitivity, and Provenance

Data adequacy and invalid states are critical in complexity and regularity estimation. Estimates can fail because support is too short, sampling is inappropriate, too few template matches or ordinal patterns occur, state space is too sparse, a recurrence threshold yields no useful recurrence structure, a scaling region is absent, a DFA fit shows crossover instead of one slope, neighbor divergence lacks a stable fit region, or missingness/interpolation dominates the result. Undefined, infinite, unstable, unidentifiable, out-of-domain, and valid-but-uncertain values must be distinguished from zero.

A worked example using synthetic signals—periodic, noisy periodic, correlated stochastic, white-noise-like, and structured irregular sequences—demonstrates why one universal complexity ranking fails. Compare Sample Entropy, Permutation Entropy, a multiscale profile, one RQA measure, one scaling/fractal descriptor, and one divergence or predictability descriptor where assumptions allow. Show sensitivity to record length, sampling, filtering or detrending, normalization, tolerance or ordinal order, recurrence threshold, embedding, scale range, and fit region. Preserve descriptor definition/version, source signal and preprocessing state, support, sampling/timing, missingness, template/embedding parameters, distance and tolerance, self-match rule, ordinal tie rule, symbolization, multiscale construction, recurrence metric/threshold/Theiler rule, scaling estimator/range, fractal estimator, neighbor and divergence settings, surrogate/null specification when used, invalid-state rules, implementation/version, uncertainty, and sensitivity findings. Conclude that reproducible computation establishes conformance to the descriptor definition, not behavioral complexity, nonlinearity, or chaos.