Linguistic Descriptors
Linguistic Descriptors are tools used in signal processing to analyze and interpret human language through structured behavioral patterns.
Linguistic Descriptors are explicit characterizations of properties of declared language-bearing evidence. These properties include lexical choice and diversity, morphological and grammatical distributions, syntactic organization, semantic content and coherence, pragmatic marking, discourse organization, repetition, repair, and other language-structured patterns. Linguistic evidence can be realized through speech, writing, signing, typing, or any other form that carries language. A transcript or symbolic encoding of such evidence is a selective representation of the original communicative event rather than the event itself. Importantly, a linguistic descriptor is not automatically a behavioral construct, psychological state, linguistic competence judgment, semantic truth, annotation, learned representation, or model output.
Meaning and Linguistic Evidence Identity
A Linguistic Descriptor is a reproducible characterization of a declared linguistic property computed from identified language-bearing evidence under explicit conventions for unitization, normalization, annotation, language, support, parameters, and interpretation. The output of a linguistic descriptor may be scalar, vector, distributional, sequence-based, relational within a linguistic structure, or structured when the output itself has stable linguistic semantics.
The following distinctions clarify related concepts:
- Language Event: The original communicative occurrence involving language use in speech, writing, signing, typing, or other forms.
- Linguistic Evidence: The language-bearing material identified from the language event, including speech signals, text, sign videos, or other representations.
- Transcript or Symbolic Encoding: A selective, symbolic representation of linguistic evidence, such as orthographic text, glosses, or phonetic transcription.
- Linguistic Annotation: The assignment of categories, labels, or relations to linguistic units under a defined scheme or ontology.
- Linguistic Descriptor: A computed measure or characterization of linguistic evidence or an explicitly declared linguistic representation.
- Feature: A measurable property or attribute derived from linguistic data, often used in modeling.
- Representation: Any encoding or formalization of linguistic evidence, including vector embeddings, parse trees, or symbolic forms.
- Model Output: An inference or prediction generated by a computational model, not a direct linguistic descriptor.
- Behavioral Construct: A scientific concept external to linguistic description, such as psychological state or social identity, which may be inferred but is not itself linguistic content.
Linguistic content differs from the physical carrier of language. Spoken language includes acoustic and vocal properties in addition to lexical, grammatical, semantic, pragmatic, and discourse structures. Written language involves orthographic and layout features. Signed language is realized through visual-manual signals with its own grammatical structure. Acoustic pitch, intensity, spectral shape, signing kinematics, handwriting geometry, or typing timing are not linguistic descriptors merely because they accompany language use.
Transcription and automatic language-conversion processes (human transcription, automatic speech recognition, sign-language annotation, optical character recognition, normalization, speaker attribution, punctuation restoration, etc.) can omit, add, merge, split, or misidentify linguistic material. Downstream descriptors therefore characterize the resulting representation under its conventions and uncertainties unless the analysis explicitly propagates uncertainty back to the original evidence.
| Object | Scientific Role | Produced or Identified By | Critical Non-Equivalence |
|---|---|---|---|
| Language Event | Original communicative occurrence | Human language users | Not fully captured by any representation or transcript |
| Transcript | Selective symbolic representation | Human transcribers, ASR, OCR, annotators | Approximation of language event, may omit or alter linguistic material |
| Token | Unit of linguistic segmentation | Tokenizers, annotators | Not always equivalent to linguistic word; depends on tokenization scheme |
| Word | Linguistic word unit | Linguistic analysis, dictionary definitions | Not always identical to tokens; language-specific definitions and morphological complexity vary |
| Lemma | Canonical base form of a word | Lemmatizers | Represents word identity abstracted from inflectional variants, not always unambiguous |
| Morpheme | Minimal meaningful linguistic unit | Morphological analysis | Different from word or token; may be overt or implicit |
| Linguistic Annotation | Categorical or relational labeling | Annotators, automatic taggers | Not equivalent to descriptors or features; assigns categories under a scheme |
| Linguistic Descriptor | Computed characterization of linguistic property | Analytical algorithms, software | Not a direct annotation, model output, or behavioral construct |
| Semantic Representation | Formal encoding of meaning | Semantic parsers, embedding models | Representation of semantic content; may be input to descriptors but is not a descriptor itself |
| Behavioral Construct | Scientific concept about behavior or psychology | Researchers, inferred from multiple data | Not a linguistic object or descriptor; distinct domain |
Linguistic Units, Segmentation, Normalization, and Annotation
Linguistic analysis employs various units of analysis, each with distinct definitions and roles. These include:
- Token: A segmented unit of text or transcription, typically defined by whitespace or other criteria.
- Orthographic Word: A word as conventionally written, often delimited by spaces or punctuation.
- Linguistic Word: A morphologically defined word unit, which may differ from orthographic words.
- Lemma: The base or dictionary form representing all inflected forms of a word.
- Stem: The core part of a word to which inflectional affixes attach.
- Morpheme: The smallest meaningful unit of language, such as root, prefix, suffix.
- Character or Grapheme: The smallest unit of written language, e.g., letters, digits, or script units.
- Phonological Symbol: Symbols representing sounds when phonological transcription is included.
- Sign or Gloss: Units representing signed language elements.
- Phrase: A group of words forming a constituent unit.
- Clause: A syntactic unit with a predicate and its arguments.
- Sentence: A complete syntactic unit expressing a statement, question, or command.
- Utterance: A spoken or signed unit bounded by speaker change or silence.
- Turn: A segment of discourse produced by one participant.
- Discourse Segment: Larger units organizing conversation or text.
The unit used by a descriptor must be explicitly declared rather than assuming universal equivalence to whitespace-delimited tokens.
Segmentation varies by language and scheme. Sentence boundaries, utterance boundaries, clauses, turns, compounds, contractions, clitics, punctuation, multiword expressions, hashtags, emojis, signed-language segmentation, and code-switch boundaries are represented differently across languages and annotation schemes. Counts normalized per sentence, clause, utterance, or token inherit these boundary definitions and must be interpreted accordingly.
Linguistic normalization involves explicit choices such as case folding, spelling normalization, punctuation handling, contraction expansion, lemmatization, stemming, diacritic treatment, numeral normalization, filler removal, stop-word removal, code normalization, and correction of recognition or transcription errors. Each transformation preserves some properties while potentially destroying others. Normalization should follow the intended linguistic property under study rather than be treated as a neutral preprocessing step.
Linguistic annotation schemes and provenance require specification of the annotation ontology or model, including part-of-speech tags, morphological features, dependency relations, constituency structures, semantic roles, named entities, discourse relations, coreference links, dialogue acts, pragmatic categories, and repair annotations. The scheme/version, language-specific extensions, annotator or model source, confidence or disagreement measures, and treatment of unassigned or ambiguous cases must be recorded.
| Decision | Descriptors It Can Affect | Failure If Hidden |
|---|---|---|
| Tokenization | Lexical frequencies, POS tagging, morphological analysis | Misinterpretation of counts, incorrect unitization |
| Sentence/Utterance Segmentation | Sentence-level lexical diversity, syntactic complexity | Inaccurate normalization, boundary-dependent measures |
| Lemmatization | Lexical diversity, semantic category prevalence | Confounded type counts, mismatched lemma-based descriptors |
| Morphological Analysis | Morphemes per word, inflectional category frequencies | Misclassification of morphological features |
| POS Tagging | POS distribution, grammatical-class composition | Misleading category proportions, invalid grammatical summaries |
| Syntactic Parsing | Dependency length, clause-type distribution | Erroneous syntactic complexity measures |
| Semantic Annotation | Semantic-role distribution, semantic coherence | Inaccurate semantic content characterization |
| Discourse/Pragmatic Annotation | Dialogue acts, coreference chains, discourse relations | Faulty discourse and pragmatic descriptor values |
| Speaker/Turn Attribution | Turn-based lexical contributions, dialogue structure | Confusion of speaker-specific patterns and interaction dynamics |
Lexical Frequency, Diversity, and Distribution
Lexical descriptors include counts and normalized frequencies over declared units such as word forms, lemmas, lexical categories, function words, content words, dictionary categories, named lexical items, or n-grams. The numerator category and denominator must be explicitly defined. Counts may be raw, per token, per utterance, per clause, or per unit time, each answering different questions and not interchangeable as normalizations.
The conventional type-token ratio (TTR) is defined as:
where N is the number of eligible lexical tokens under the declared tokenization and filtering rules with N > 0, V is the number of distinct eligible types under the declared type identity (e.g., surface form or lemma), and TTR is the type-token ratio. TTR depends strongly on sample length and on the definition of tokens and types; values from substantially different support lengths should not be directly compared as measures of lexical diversity without further justification.
Lexical diversity is a family of measures rather than a single statistic. Ordinary TTR is conceptually compared with moving-average TTR, root/log-corrected ratios, vocd-D-like approaches, HD-D-like approaches, and MTLD-like measures. Different methods address the influence of text length through different constructions. Parameters such as window length, threshold, sampling, randomization, token/type definition, and minimum support requirements must be explicit. No measure is universally length-invariant or superior.
Lexical density and grammatical-class composition are described through declared categories such as nouns, lexical verbs, adjectives, adverbs, function words, pronouns, auxiliaries, or language-specific classes. The annotation scheme and denominator must be stated. Content/function distinctions are linguistic analyses that vary across languages, and greater lexical density should not be interpreted automatically as better language, greater competence, or greater cognitive sophistication.
Lexical frequency-profile, rarity, familiarity, register, concreteness, imageability, affective-lexicon, or other lexicon-referenced descriptors require explicit reporting of the reference resource, language, corpus population, scale direction, coverage, version, and out-of-vocabulary policy. A word can be rare in one reference corpus and common in another. Lexicon categories operationalize lexical properties rather than direct behavioral or psychological states.
Morphological and Grammatical Descriptors
Morphological descriptors characterize properties such as morphemes per word, inflectional or derivational category frequencies, tense/aspect/mood distributions, person/number/gender/case marking where encoded, affix or clitic distributions, morphological-type diversity, and selected paradigm-use patterns. Analyses require language-specific morphological knowledge. Absence of an overt marker does not necessarily indicate absence of a grammatical function when linguistic realization varies.
Part-of-speech and morphosyntactic-feature distributions describe grammatical category use under a declared tagset. Counts, proportions, entropy-like diversity, transition patterns, and co-occurrence summaries are legitimate descriptors but inherit tagging errors, tokenization choices, language-specific categories, and annotation conventions. Universal tag inventories improve comparability but do not erase real grammatical differences among languages.
Grammatical-construction descriptors cover clause-type distribution, subordination/coordination prevalence, negation constructions, question forms, passive or voice constructions where applicable, agreement patterns, argument-structure patterns, or other declared constructions. Definitions must be language-appropriate. Observed construction frequency is distinct from grammatical ability, preference, or communicative intention.
| Descriptor | Required Linguistic Analysis | Typical Denominator | Interpretive Caution |
|---|---|---|---|
| Morphological Feature Frequency | Morphological analysis | Number of tokens or words | Tagger errors, language-specific categories, incomplete analysis |
| Morphemes per Word | Morphological segmentation | Word count | Varies with morphological complexity and tokenization decisions |
| POS Distribution | POS tagging | Token count | Tagset dependence, annotation errors, language differences |
| Function/Content Composition | POS tagging, lexical categorization | Token or word count | Variable definitions across languages; not universal cognitive indices |
| Clause-Type Distribution | Syntactic parsing or annotation | Clause count | Parsing scheme sensitivity and definition of clause types |
| Subordination/Coordination | Syntactic parsing or annotation | Clause or sentence count | Annotation scheme and syntactic theory differences |
| Construction Frequency | Construction identification | Total constructions or tokens | Requires explicit construction definitions; not necessarily ability |
Syntactic Structure Descriptors
Syntactic complexity is a multidimensional family of descriptors rather than a single universal scalar. Representative descriptors include sentence or clause length, clauses per sentence, subordination, coordination, constituent depth, dependency length, branching factor, dependency-relation distribution, noun-phrase or verb-phrase elaboration, and structural diversity. Different descriptors capture distinct aspects of syntax and may vary independently across languages, genres, tasks, and support lengths.
Mean dependency distance (MDD) is defined as:
where D is the declared set of eligible dependency arcs in the analyzed syntactic structure, K is the number of arcs in D with K > 0, j is an eligible dependent token, p_j is its linear token position under the declared tokenization, h(j) is the syntactic head of token j, p_{h(j)} is that head's token position, and MDD is the mean absolute linear dependency distance. Punctuation handling, root relations, multiword tokens, coordination analyses, language word order, tokenization, and parser scheme can alter the value. Longer dependency distance does not directly measure cognitive load for an individual.
Tree- and clause-based syntactic descriptors include dependency-tree depth, constituency depth (when constituency analysis exists), branching factor, clause length, phrase length, number of dependents, embedded-clause depth, and syntactic-pattern diversity. The parse formalism and treatment of punctuation, coordination, ellipsis, fragments, and incomplete utterances must be specified. Tree depth under one grammar is not numerically interchangeable with tree depth under another.
Parser and annotation uncertainty must be recognized. Automatically parsed transcripts may contain correlated errors from recognition, punctuation restoration, segmentation, tagging, and parsing. Descriptor pipelines should preserve parser/model/version and confidence or error information when available. Sensitivity to plausible alternative parses should be considered, especially for short, disfluent, code-switched, conversational, or nonstandard language.
Semantic Content and Coherence Descriptors
Semantic-category descriptors derive from declared dictionaries, ontologies, semantic-role schemes, lexical resources, or manually defined categories. Outputs include category prevalence, semantic-role distributions, entity or referent categories, concreteness or affective lexical dimensions, and domain-semantic content profiles. Resource coverage, language, sense or lemma mapping, category overlap, negation/context policy, and out-of-resource handling must be stated.
Semantic similarity and coherence descriptors quantify relations among words, utterances, sentences, turns, or discourse segments under a declared semantic representation. Outputs include adjacent-segment similarity, global-to-local coherence, semantic repetition, semantic drift, reference-to-topic similarity, or dispersion within a semantic space. Representation/model/version, pooling, normalization, similarity or distance function, unitization, and context window are required. High vector similarity indicates similarity under that representation, not proof of equivalent human meaning.
Topic, theme, entity, semantic-role, and proposition-like descriptors require caution. Manually defined themes, lexicon categories, probabilistic topic models, classifiers, and language-model-derived categories produce different objects. A learned topic distribution or classifier probability is a model output or representation unless a stable descriptor definition explicitly assigns interpretable semantics. Predictive utility alone does not convert latent dimensions into linguistic facts.
| Descriptor | Required Representation | Property Characterized | Main Interpretation Risk |
|---|---|---|---|
| Dictionary/Semantic Category Prevalence | Lexicon or ontology-based mapping | Category frequency or coverage | Resource coverage limits, ambiguity, context sensitivity |
| Semantic-Role Distribution | Semantic role annotation or parsing | Semantic argument structure | Annotation scheme dependence, role granularity |
| Entity/Referent Descriptor | Coreference or entity linking | Referent types and properties | Linking errors, ambiguous references |
| Adjacent Semantic Similarity | Vector embeddings or semantic models | Local semantic coherence | Representation-dependent similarity; not direct equivalence |
| Global Coherence | Semantic representations with pooling | Overall discourse coherence | Dependent on model, window size, and normalization |
| Semantic Drift | Semantic trajectory over segments | Topic or meaning change over time | Sensitive to segmentation and representation choice |
| Topic/Theme Descriptor | Topic models, classifiers | Thematic content distribution | Model assumptions, interpretability of topics |
Pragmatic, Discourse, and Repair Descriptors
Discourse-organization and cohesion descriptors include discourse-connective distributions, referential overlap, coreference-chain properties, lexical cohesion, entity continuity, discourse-relation frequencies, topic continuity, and segment linkage. The discourse segmentation and annotation or representation scheme used must be specified. Cohesion is a property of linguistic organization under a defined framework and should not be equated automatically with coherence, quality, truthfulness, or communicative success.
Pragmatic and interaction-function descriptors cover question forms, directives, acknowledgments, hedges, stance markers, politeness markers, discourse markers, dialogue-act distributions, deixis, reported speech, quotation, and other operationally defined language-use categories. These characterize linguistic choices in context but do not directly reveal intention, certainty, sincerity, emotion, dominance, deception, or social relationships.
Disfluency and repair descriptors encompass filled pauses (when linguistically represented), repetitions, false starts, abandoned constructions, substitutions, self-corrections, repair initiations, and repair completions. Transcription convention and event identity must be recorded. Linguistic occurrence of repair is distinct from acoustic timing or vocal realization. Disfluency frequency should not be treated as a direct diagnostic or psychological measure.
Turn- and utterance-level descriptors include utterance-function distribution, question-response structure, lexical contribution per turn, turn-internal syntactic form, response semantic relevance, and discourse-role categories when conversation is the linguistic object. Speaker and turn identity must be preserved. Linguistic description of conversational contributions is distinct from measures of cross-person synchrony, coupling, coordination, or causal influence.
Sequence, Repetition, and Composite Linguistic Descriptors
Linguistic sequence descriptors are based on ordered tokens, lemmas, POS tags, morphological categories, discourse acts, or other declared linguistic symbols. Representative properties include n-gram frequencies, repetition distance, recurrence of lexical or grammatical patterns, transition distributions, alternation patterns, and sequential diversity. Unordered bag-of-items representations preserve counts but destroy order, adjacency, discourse progression, and syntactic relationships.
Distributional and information-theoretic summaries of linguistic categories, such as entropy of token, lemma, POS, construction, or dialogue-act distributions, require explicit linguistic alphabet, support, estimator, normalization, and ordering assumptions. Static category entropy differs from sequence entropy, conditional entropy, language-model perplexity, and general complexity descriptors. A high lexical-category entropy is not a universal measure of linguistic sophistication.
Composite linguistic indices include readability-like scores, syntactic-complexity composites, lexical sophistication composites, cohesion composites, or style indices formed as convention-dependent combinations of lower-level descriptors. Component definitions, weights, language, genre, intended population, calibration source, and version must be explicit. Composite scores can be useful operational summaries but should not be treated as universal measures of linguistic difficulty, competence, intelligence, cognitive load, or behavioral quality.
| Descriptor | Ordering Preserved? | Required Definition | Overclaim to Avoid |
|---|---|---|---|
| N-Gram Frequency | Yes | Unitization, normalization, n value | Treating counts as context-free or frequency-only |
| Lexical Repetition | Yes | Token or lemma identity, distance metric | Equating repetition with communicative quality |
| POS/Construction Transition | Yes | Tagset, unitization | Assuming transitions reflect universal syntax |
| Static Category Entropy | No | Alphabet, support definition | Inferring sophistication from entropy alone |
| Conditional/Sequential Category Descriptor | Yes | Model, order, estimator, normalization | Treating perplexity as cognitive difficulty |
| Readability-Like Composite | No | Component descriptors, weights, calibration | Universal measure of intelligence or difficulty |
| Linguistic Style Composite | No | Component definitions, domain specificity | Attributing style to personality or competence |
Adequacy, Context, Sensitivity, and Provenance
An integrated provenance audit contrasts two language samples matched on task and language but differing in length, alongside an automatically transcribed or parsed variant containing realistic representation errors.
Consider:
- Sample A: Human-transcribed spoken narrative, length 1000 tokens.
- Sample B: Human-transcribed narrative on the same topic and task, length 500 tokens.
- Sample C: Automatic speech recognition (ASR) transcript of Sample A, including segmentation and recognition errors.
Lexical Counts and Normalization:
-
Raw lexical counts in Sample A exceed Sample B due to length.
-
Normalized frequencies per 100 tokens reveal comparable category proportions.
-
TTR (type-token ratio) computed with tokenization and filtering:
Sample A:
with V=450, N=1000, TTR=0.45.Sample B:
V=280, N=500, TTR=0.56. -
A length-robust lexical diversity measure (e.g., MTLD) accounts for sample length differences, showing more comparable diversity.
Morphological/POS Descriptor:
- POS distribution normalized by token count shows consistent proportions in nouns, verbs, function words across Samples A and B.
- ASR errors in Sample C cause misclassifications, distorting POS proportions.
Syntactic Descriptor:
- Mean dependency distance (MDD) calculated on dependency parses reflects sentence complexity.
- Parsing errors in Sample C increase variability.
- Differences in tokenization and punctuation handling influence MDD values.
Semantic Descriptor:
- Semantic category prevalence measured using a lexicon resource with known coverage.
- Sample C’s automatic transcript misses some lexical items, reducing coverage.
- Category prevalence normalized by token count.
Discourse/Pragmatic Descriptor:
- Dialogue-act distribution or discourse-connective frequency normalized per utterance.
- Speaker turn identity preserved in Samples A and B.
- ASR transcript conflates speaker turns, obfuscating pragmatic patterns.
Sequence/Repetition Descriptor:
- N-gram frequencies and repetition distances computed on tokens and lemmas.
- Errors in Sample C reduce observed repetition.
- Normalization by token count and explicit unitization affect comparability.
Throughout, documented parameters include:
- Descriptor definition and version
- Language (and dialect/register if known)
- Source medium (spoken, written, signed)
- Speaker/author identity and turn attribution
- Transcript or representation version
- Tokenization and segmentation scheme
- Normalization steps applied
- Annotation scheme/model/version
- Units and denominators used
- Lexical diversity parameters
- Parser formalism and version
- Semantic resource or embedding model used
- Discourse/pragmatic ontology
- Missing and unknown category handling
- Confidence or disagreement metrics
- Implementation version
- Uncertainty and sensitivity analyses
Interpretation emphasizes that reproducible linguistic computation establishes conformance to descriptor definitions and conventions. It does not provide direct evidence of emotion, intention, deception, personality, intelligence, diagnosis, identity, competence, social status, causal mechanisms, or universal behavioral meaning. Contextual matching and measurement invariance are necessary for valid comparisons, and observed linguistic differences must not be automatically equated with behavioral or psychological differences.