✦ For everyone, free.

Practical knowledge for real and everyday life

Home

Linguistic Descriptors

Linguistic Descriptors are tools used in signal processing to analyze and interpret human language through structured behavioral patterns.

Linguistic Descriptors are explicit characterizations of properties of declared language-bearing evidence. These properties include lexical choice and diversity, morphological and grammatical distributions, syntactic organization, semantic content and coherence, pragmatic marking, discourse organization, repetition, repair, and other language-structured patterns. Linguistic evidence can be realized through speech, writing, signing, typing, or any other form that carries language. A transcript or symbolic encoding of such evidence is a selective representation of the original communicative event rather than the event itself. Importantly, a linguistic descriptor is not automatically a behavioral construct, psychological state, linguistic competence judgment, semantic truth, annotation, learned representation, or model output.


Meaning and Linguistic Evidence Identity

A Linguistic Descriptor is a reproducible characterization of a declared linguistic property computed from identified language-bearing evidence under explicit conventions for unitization, normalization, annotation, language, support, parameters, and interpretation. The output of a linguistic descriptor may be scalar, vector, distributional, sequence-based, relational within a linguistic structure, or structured when the output itself has stable linguistic semantics.

The following distinctions clarify related concepts:

  • Language Event: The original communicative occurrence involving language use in speech, writing, signing, typing, or other forms.
  • Linguistic Evidence: The language-bearing material identified from the language event, including speech signals, text, sign videos, or other representations.
  • Transcript or Symbolic Encoding: A selective, symbolic representation of linguistic evidence, such as orthographic text, glosses, or phonetic transcription.
  • Linguistic Annotation: The assignment of categories, labels, or relations to linguistic units under a defined scheme or ontology.
  • Linguistic Descriptor: A computed measure or characterization of linguistic evidence or an explicitly declared linguistic representation.
  • Feature: A measurable property or attribute derived from linguistic data, often used in modeling.
  • Representation: Any encoding or formalization of linguistic evidence, including vector embeddings, parse trees, or symbolic forms.
  • Model Output: An inference or prediction generated by a computational model, not a direct linguistic descriptor.
  • Behavioral Construct: A scientific concept external to linguistic description, such as psychological state or social identity, which may be inferred but is not itself linguistic content.

Linguistic content differs from the physical carrier of language. Spoken language includes acoustic and vocal properties in addition to lexical, grammatical, semantic, pragmatic, and discourse structures. Written language involves orthographic and layout features. Signed language is realized through visual-manual signals with its own grammatical structure. Acoustic pitch, intensity, spectral shape, signing kinematics, handwriting geometry, or typing timing are not linguistic descriptors merely because they accompany language use.

Transcription and automatic language-conversion processes (human transcription, automatic speech recognition, sign-language annotation, optical character recognition, normalization, speaker attribution, punctuation restoration, etc.) can omit, add, merge, split, or misidentify linguistic material. Downstream descriptors therefore characterize the resulting representation under its conventions and uncertainties unless the analysis explicitly propagates uncertainty back to the original evidence.

ObjectScientific RoleProduced or Identified ByCritical Non-Equivalence
Language EventOriginal communicative occurrenceHuman language usersNot fully captured by any representation or transcript
TranscriptSelective symbolic representationHuman transcribers, ASR, OCR, annotatorsApproximation of language event, may omit or alter linguistic material
TokenUnit of linguistic segmentationTokenizers, annotatorsNot always equivalent to linguistic word; depends on tokenization scheme
WordLinguistic word unitLinguistic analysis, dictionary definitionsNot always identical to tokens; language-specific definitions and morphological complexity vary
LemmaCanonical base form of a wordLemmatizersRepresents word identity abstracted from inflectional variants, not always unambiguous
MorphemeMinimal meaningful linguistic unitMorphological analysisDifferent from word or token; may be overt or implicit
Linguistic AnnotationCategorical or relational labelingAnnotators, automatic taggersNot equivalent to descriptors or features; assigns categories under a scheme
Linguistic DescriptorComputed characterization of linguistic propertyAnalytical algorithms, softwareNot a direct annotation, model output, or behavioral construct
Semantic RepresentationFormal encoding of meaningSemantic parsers, embedding modelsRepresentation of semantic content; may be input to descriptors but is not a descriptor itself
Behavioral ConstructScientific concept about behavior or psychologyResearchers, inferred from multiple dataNot a linguistic object or descriptor; distinct domain

Linguistic Units, Segmentation, Normalization, and Annotation

Linguistic analysis employs various units of analysis, each with distinct definitions and roles. These include:

  • Token: A segmented unit of text or transcription, typically defined by whitespace or other criteria.
  • Orthographic Word: A word as conventionally written, often delimited by spaces or punctuation.
  • Linguistic Word: A morphologically defined word unit, which may differ from orthographic words.
  • Lemma: The base or dictionary form representing all inflected forms of a word.
  • Stem: The core part of a word to which inflectional affixes attach.
  • Morpheme: The smallest meaningful unit of language, such as root, prefix, suffix.
  • Character or Grapheme: The smallest unit of written language, e.g., letters, digits, or script units.
  • Phonological Symbol: Symbols representing sounds when phonological transcription is included.
  • Sign or Gloss: Units representing signed language elements.
  • Phrase: A group of words forming a constituent unit.
  • Clause: A syntactic unit with a predicate and its arguments.
  • Sentence: A complete syntactic unit expressing a statement, question, or command.
  • Utterance: A spoken or signed unit bounded by speaker change or silence.
  • Turn: A segment of discourse produced by one participant.
  • Discourse Segment: Larger units organizing conversation or text.

The unit used by a descriptor must be explicitly declared rather than assuming universal equivalence to whitespace-delimited tokens.

Segmentation varies by language and scheme. Sentence boundaries, utterance boundaries, clauses, turns, compounds, contractions, clitics, punctuation, multiword expressions, hashtags, emojis, signed-language segmentation, and code-switch boundaries are represented differently across languages and annotation schemes. Counts normalized per sentence, clause, utterance, or token inherit these boundary definitions and must be interpreted accordingly.

Linguistic normalization involves explicit choices such as case folding, spelling normalization, punctuation handling, contraction expansion, lemmatization, stemming, diacritic treatment, numeral normalization, filler removal, stop-word removal, code normalization, and correction of recognition or transcription errors. Each transformation preserves some properties while potentially destroying others. Normalization should follow the intended linguistic property under study rather than be treated as a neutral preprocessing step.

Linguistic annotation schemes and provenance require specification of the annotation ontology or model, including part-of-speech tags, morphological features, dependency relations, constituency structures, semantic roles, named entities, discourse relations, coreference links, dialogue acts, pragmatic categories, and repair annotations. The scheme/version, language-specific extensions, annotator or model source, confidence or disagreement measures, and treatment of unassigned or ambiguous cases must be recorded.

DecisionDescriptors It Can AffectFailure If Hidden
TokenizationLexical frequencies, POS tagging, morphological analysisMisinterpretation of counts, incorrect unitization
Sentence/Utterance SegmentationSentence-level lexical diversity, syntactic complexityInaccurate normalization, boundary-dependent measures
LemmatizationLexical diversity, semantic category prevalenceConfounded type counts, mismatched lemma-based descriptors
Morphological AnalysisMorphemes per word, inflectional category frequenciesMisclassification of morphological features
POS TaggingPOS distribution, grammatical-class compositionMisleading category proportions, invalid grammatical summaries
Syntactic ParsingDependency length, clause-type distributionErroneous syntactic complexity measures
Semantic AnnotationSemantic-role distribution, semantic coherenceInaccurate semantic content characterization
Discourse/Pragmatic AnnotationDialogue acts, coreference chains, discourse relationsFaulty discourse and pragmatic descriptor values
Speaker/Turn AttributionTurn-based lexical contributions, dialogue structureConfusion of speaker-specific patterns and interaction dynamics

Lexical Frequency, Diversity, and Distribution

Lexical descriptors include counts and normalized frequencies over declared units such as word forms, lemmas, lexical categories, function words, content words, dictionary categories, named lexical items, or n-grams. The numerator category and denominator must be explicitly defined. Counts may be raw, per token, per utterance, per clause, or per unit time, each answering different questions and not interchangeable as normalizations.

The conventional type-token ratio (TTR) is defined as:

TTR = V N

where N is the number of eligible lexical tokens under the declared tokenization and filtering rules with N > 0, V is the number of distinct eligible types under the declared type identity (e.g., surface form or lemma), and TTR is the type-token ratio. TTR depends strongly on sample length and on the definition of tokens and types; values from substantially different support lengths should not be directly compared as measures of lexical diversity without further justification.

Lexical diversity is a family of measures rather than a single statistic. Ordinary TTR is conceptually compared with moving-average TTR, root/log-corrected ratios, vocd-D-like approaches, HD-D-like approaches, and MTLD-like measures. Different methods address the influence of text length through different constructions. Parameters such as window length, threshold, sampling, randomization, token/type definition, and minimum support requirements must be explicit. No measure is universally length-invariant or superior.

Lexical density and grammatical-class composition are described through declared categories such as nouns, lexical verbs, adjectives, adverbs, function words, pronouns, auxiliaries, or language-specific classes. The annotation scheme and denominator must be stated. Content/function distinctions are linguistic analyses that vary across languages, and greater lexical density should not be interpreted automatically as better language, greater competence, or greater cognitive sophistication.

Lexical frequency-profile, rarity, familiarity, register, concreteness, imageability, affective-lexicon, or other lexicon-referenced descriptors require explicit reporting of the reference resource, language, corpus population, scale direction, coverage, version, and out-of-vocabulary policy. A word can be rare in one reference corpus and common in another. Lexicon categories operationalize lexical properties rather than direct behavioral or psychological states.


Morphological and Grammatical Descriptors

Morphological descriptors characterize properties such as morphemes per word, inflectional or derivational category frequencies, tense/aspect/mood distributions, person/number/gender/case marking where encoded, affix or clitic distributions, morphological-type diversity, and selected paradigm-use patterns. Analyses require language-specific morphological knowledge. Absence of an overt marker does not necessarily indicate absence of a grammatical function when linguistic realization varies.

Part-of-speech and morphosyntactic-feature distributions describe grammatical category use under a declared tagset. Counts, proportions, entropy-like diversity, transition patterns, and co-occurrence summaries are legitimate descriptors but inherit tagging errors, tokenization choices, language-specific categories, and annotation conventions. Universal tag inventories improve comparability but do not erase real grammatical differences among languages.

Grammatical-construction descriptors cover clause-type distribution, subordination/coordination prevalence, negation constructions, question forms, passive or voice constructions where applicable, agreement patterns, argument-structure patterns, or other declared constructions. Definitions must be language-appropriate. Observed construction frequency is distinct from grammatical ability, preference, or communicative intention.

DescriptorRequired Linguistic AnalysisTypical DenominatorInterpretive Caution
Morphological Feature FrequencyMorphological analysisNumber of tokens or wordsTagger errors, language-specific categories, incomplete analysis
Morphemes per WordMorphological segmentationWord countVaries with morphological complexity and tokenization decisions
POS DistributionPOS taggingToken countTagset dependence, annotation errors, language differences
Function/Content CompositionPOS tagging, lexical categorizationToken or word countVariable definitions across languages; not universal cognitive indices
Clause-Type DistributionSyntactic parsing or annotationClause countParsing scheme sensitivity and definition of clause types
Subordination/CoordinationSyntactic parsing or annotationClause or sentence countAnnotation scheme and syntactic theory differences
Construction FrequencyConstruction identificationTotal constructions or tokensRequires explicit construction definitions; not necessarily ability

Syntactic Structure Descriptors

Syntactic complexity is a multidimensional family of descriptors rather than a single universal scalar. Representative descriptors include sentence or clause length, clauses per sentence, subordination, coordination, constituent depth, dependency length, branching factor, dependency-relation distribution, noun-phrase or verb-phrase elaboration, and structural diversity. Different descriptors capture distinct aspects of syntax and may vary independently across languages, genres, tasks, and support lengths.

Mean dependency distance (MDD) is defined as:

MDD = 1 K j D | p j p h ( j ) |

where D is the declared set of eligible dependency arcs in the analyzed syntactic structure, K is the number of arcs in D with K > 0, j is an eligible dependent token, p_j is its linear token position under the declared tokenization, h(j) is the syntactic head of token j, p_{h(j)} is that head's token position, and MDD is the mean absolute linear dependency distance. Punctuation handling, root relations, multiword tokens, coordination analyses, language word order, tokenization, and parser scheme can alter the value. Longer dependency distance does not directly measure cognitive load for an individual.

Tree- and clause-based syntactic descriptors include dependency-tree depth, constituency depth (when constituency analysis exists), branching factor, clause length, phrase length, number of dependents, embedded-clause depth, and syntactic-pattern diversity. The parse formalism and treatment of punctuation, coordination, ellipsis, fragments, and incomplete utterances must be specified. Tree depth under one grammar is not numerically interchangeable with tree depth under another.

Parser and annotation uncertainty must be recognized. Automatically parsed transcripts may contain correlated errors from recognition, punctuation restoration, segmentation, tagging, and parsing. Descriptor pipelines should preserve parser/model/version and confidence or error information when available. Sensitivity to plausible alternative parses should be considered, especially for short, disfluent, code-switched, conversational, or nonstandard language.


Semantic Content and Coherence Descriptors

Semantic-category descriptors derive from declared dictionaries, ontologies, semantic-role schemes, lexical resources, or manually defined categories. Outputs include category prevalence, semantic-role distributions, entity or referent categories, concreteness or affective lexical dimensions, and domain-semantic content profiles. Resource coverage, language, sense or lemma mapping, category overlap, negation/context policy, and out-of-resource handling must be stated.

Semantic similarity and coherence descriptors quantify relations among words, utterances, sentences, turns, or discourse segments under a declared semantic representation. Outputs include adjacent-segment similarity, global-to-local coherence, semantic repetition, semantic drift, reference-to-topic similarity, or dispersion within a semantic space. Representation/model/version, pooling, normalization, similarity or distance function, unitization, and context window are required. High vector similarity indicates similarity under that representation, not proof of equivalent human meaning.

Topic, theme, entity, semantic-role, and proposition-like descriptors require caution. Manually defined themes, lexicon categories, probabilistic topic models, classifiers, and language-model-derived categories produce different objects. A learned topic distribution or classifier probability is a model output or representation unless a stable descriptor definition explicitly assigns interpretable semantics. Predictive utility alone does not convert latent dimensions into linguistic facts.

DescriptorRequired RepresentationProperty CharacterizedMain Interpretation Risk
Dictionary/Semantic Category PrevalenceLexicon or ontology-based mappingCategory frequency or coverageResource coverage limits, ambiguity, context sensitivity
Semantic-Role DistributionSemantic role annotation or parsingSemantic argument structureAnnotation scheme dependence, role granularity
Entity/Referent DescriptorCoreference or entity linkingReferent types and propertiesLinking errors, ambiguous references
Adjacent Semantic SimilarityVector embeddings or semantic modelsLocal semantic coherenceRepresentation-dependent similarity; not direct equivalence
Global CoherenceSemantic representations with poolingOverall discourse coherenceDependent on model, window size, and normalization
Semantic DriftSemantic trajectory over segmentsTopic or meaning change over timeSensitive to segmentation and representation choice
Topic/Theme DescriptorTopic models, classifiersThematic content distributionModel assumptions, interpretability of topics

Pragmatic, Discourse, and Repair Descriptors

Discourse-organization and cohesion descriptors include discourse-connective distributions, referential overlap, coreference-chain properties, lexical cohesion, entity continuity, discourse-relation frequencies, topic continuity, and segment linkage. The discourse segmentation and annotation or representation scheme used must be specified. Cohesion is a property of linguistic organization under a defined framework and should not be equated automatically with coherence, quality, truthfulness, or communicative success.

Pragmatic and interaction-function descriptors cover question forms, directives, acknowledgments, hedges, stance markers, politeness markers, discourse markers, dialogue-act distributions, deixis, reported speech, quotation, and other operationally defined language-use categories. These characterize linguistic choices in context but do not directly reveal intention, certainty, sincerity, emotion, dominance, deception, or social relationships.

Disfluency and repair descriptors encompass filled pauses (when linguistically represented), repetitions, false starts, abandoned constructions, substitutions, self-corrections, repair initiations, and repair completions. Transcription convention and event identity must be recorded. Linguistic occurrence of repair is distinct from acoustic timing or vocal realization. Disfluency frequency should not be treated as a direct diagnostic or psychological measure.

Turn- and utterance-level descriptors include utterance-function distribution, question-response structure, lexical contribution per turn, turn-internal syntactic form, response semantic relevance, and discourse-role categories when conversation is the linguistic object. Speaker and turn identity must be preserved. Linguistic description of conversational contributions is distinct from measures of cross-person synchrony, coupling, coordination, or causal influence.


Sequence, Repetition, and Composite Linguistic Descriptors

Linguistic sequence descriptors are based on ordered tokens, lemmas, POS tags, morphological categories, discourse acts, or other declared linguistic symbols. Representative properties include n-gram frequencies, repetition distance, recurrence of lexical or grammatical patterns, transition distributions, alternation patterns, and sequential diversity. Unordered bag-of-items representations preserve counts but destroy order, adjacency, discourse progression, and syntactic relationships.

Distributional and information-theoretic summaries of linguistic categories, such as entropy of token, lemma, POS, construction, or dialogue-act distributions, require explicit linguistic alphabet, support, estimator, normalization, and ordering assumptions. Static category entropy differs from sequence entropy, conditional entropy, language-model perplexity, and general complexity descriptors. A high lexical-category entropy is not a universal measure of linguistic sophistication.

Composite linguistic indices include readability-like scores, syntactic-complexity composites, lexical sophistication composites, cohesion composites, or style indices formed as convention-dependent combinations of lower-level descriptors. Component definitions, weights, language, genre, intended population, calibration source, and version must be explicit. Composite scores can be useful operational summaries but should not be treated as universal measures of linguistic difficulty, competence, intelligence, cognitive load, or behavioral quality.

DescriptorOrdering Preserved?Required DefinitionOverclaim to Avoid
N-Gram FrequencyYesUnitization, normalization, n valueTreating counts as context-free or frequency-only
Lexical RepetitionYesToken or lemma identity, distance metricEquating repetition with communicative quality
POS/Construction TransitionYesTagset, unitizationAssuming transitions reflect universal syntax
Static Category EntropyNoAlphabet, support definitionInferring sophistication from entropy alone
Conditional/Sequential Category DescriptorYesModel, order, estimator, normalizationTreating perplexity as cognitive difficulty
Readability-Like CompositeNoComponent descriptors, weights, calibrationUniversal measure of intelligence or difficulty
Linguistic Style CompositeNoComponent definitions, domain specificityAttributing style to personality or competence

Adequacy, Context, Sensitivity, and Provenance

An integrated provenance audit contrasts two language samples matched on task and language but differing in length, alongside an automatically transcribed or parsed variant containing realistic representation errors.

Consider:

  • Sample A: Human-transcribed spoken narrative, length 1000 tokens.
  • Sample B: Human-transcribed narrative on the same topic and task, length 500 tokens.
  • Sample C: Automatic speech recognition (ASR) transcript of Sample A, including segmentation and recognition errors.

Lexical Counts and Normalization:

  • Raw lexical counts in Sample A exceed Sample B due to length.

  • Normalized frequencies per 100 tokens reveal comparable category proportions.

  • TTR (type-token ratio) computed with tokenization and filtering:

    Sample A:

    TTR=VN with V=450, N=1000, TTR=0.45.

    Sample B:
    V=280, N=500, TTR=0.56.

  • A length-robust lexical diversity measure (e.g., MTLD) accounts for sample length differences, showing more comparable diversity.

Morphological/POS Descriptor:

  • POS distribution normalized by token count shows consistent proportions in nouns, verbs, function words across Samples A and B.
  • ASR errors in Sample C cause misclassifications, distorting POS proportions.

Syntactic Descriptor:

  • Mean dependency distance (MDD) calculated on dependency parses reflects sentence complexity.
  • Parsing errors in Sample C increase variability.
  • Differences in tokenization and punctuation handling influence MDD values.

Semantic Descriptor:

  • Semantic category prevalence measured using a lexicon resource with known coverage.
  • Sample C’s automatic transcript misses some lexical items, reducing coverage.
  • Category prevalence normalized by token count.

Discourse/Pragmatic Descriptor:

  • Dialogue-act distribution or discourse-connective frequency normalized per utterance.
  • Speaker turn identity preserved in Samples A and B.
  • ASR transcript conflates speaker turns, obfuscating pragmatic patterns.

Sequence/Repetition Descriptor:

  • N-gram frequencies and repetition distances computed on tokens and lemmas.
  • Errors in Sample C reduce observed repetition.
  • Normalization by token count and explicit unitization affect comparability.

Throughout, documented parameters include:

  • Descriptor definition and version
  • Language (and dialect/register if known)
  • Source medium (spoken, written, signed)
  • Speaker/author identity and turn attribution
  • Transcript or representation version
  • Tokenization and segmentation scheme
  • Normalization steps applied
  • Annotation scheme/model/version
  • Units and denominators used
  • Lexical diversity parameters
  • Parser formalism and version
  • Semantic resource or embedding model used
  • Discourse/pragmatic ontology
  • Missing and unknown category handling
  • Confidence or disagreement metrics
  • Implementation version
  • Uncertainty and sensitivity analyses

Interpretation emphasizes that reproducible linguistic computation establishes conformance to descriptor definitions and conventions. It does not provide direct evidence of emotion, intention, deception, personality, intelligence, diagnosis, identity, competence, social status, causal mechanisms, or universal behavioral meaning. Contextual matching and measurement invariance are necessary for valid comparisons, and observed linguistic differences must not be automatically equated with behavioral or psychological differences.