CORTEXA
← Browse
openalexOpen MINDCited by 0

The Sociolinguistics of Machine Identity: LLM Personality and Ideology Propagation

Guangni Li

Do large language models (LLMs) possess a measurable "personality," and how do the linguistic properties of training corpora shape their cognitive style and downstream reasoning? This paper approaches these questions from a sociolinguistic perspective on machine "identity." This manuscript is positioned explicitly as a conceptual perspective paper: it does not present original experimental data, nor does it constitute a systematic review with defined search strategies or inclusion criteria. Instead, it synthesises key published findings into two original theoretical frameworks intended to guide future empirical and engineering work. We examine evidence that LLMs exhibit reliable, valid Big Five personality traits—particularly in large, instruction-tuned models—and that continued pre-training on domain-specific corpora may shape those traits through measurable linguistic features: imperative ratio, type-token ratio (TTR), and syntactic complexity. We then analyse how standard-language ideology embedded in training corpora is amplified in model outputs, disadvantaging dialect and minority-language communities. Building on these findings, we propose two conceptual contributions: (1) a Personality Engineering (PE) framework for targeted continued pre-training to cultivate task-appropriate cognitive profiles, and (2) a Language Ideology Propagation Model (LIPM) mapping the pipeline from corpus composition to societal impact. Both frameworks are explicitly conceptual and require empirical validation before they can function as operational guidelines. Their value lies in structuring future research and providing a shared vocabulary for cross-disciplinary collaboration, with direct implications for responsible AI deployment.

View free PDFSource page

Related papers

openalexOpen MIND2026-07-23

The Volatility Channel (Financial Branch)

Jean-Pierre Bronsard

🔗 Reproducible code: github.com/jpbronsard/syntonic-portfolio v 3.0 V2.0 measured one channel of financial adaptation: the return channel, where \(\tau^\star=1/\sqrt{2}\) is the structural signature of the random-walk limit. Markets have a second channel with the same structure…

Source page
openalexOpen MIND2026-07-24

Creepybits/software-defined-stochastic-inference-engine: SDSIE v1.0.1 - Initial Specification & Blueprint

Creepybits

Software-Defined Stochastic Inference Engine (SDSIE) Initial public release and technical research specification for the Software-Defined Stochastic Inference Engine (SDSIE). Overview This specification outlines an adaptive virtual runtime architecture designed to execute generat…

View free PDFSource page
openalexOpen MIND2026-07-23

Friction-Guided Inference: Calibrating Correction Strategies and Abstention from Logprob Signals

Tomas Pødenphant Lund

Large language models frequently possess the knowledge needed to answer a question correctly yet commit to the wrong response. This paper presents friction-guided inference, a calibrated inference-time pipeline that uses the model's own logprob distribution — available at zero co…

View free PDFSource page
openalexOpen MIND2026-07-25

RivanDioPerdinan/Prediction-Microemulsion-Phase-Classes-Formulation-Variables-Using-ML-And-FI: Microemulsion Phase-Class Prediction Workflow v2.0.0

Dio

Final reproducible workflow for the manuscript "Prediction of Microemulsion Phase Classes from Formulation Variables Using Machine Learning and Feature Importance". This release contains the raw dataset, cleaned dataset, data dictionary, Python workflow, package requirements, gen…

View free PDFSource page
openalexOpen MIND2026-07-30

Retrieval Augmented Generation Using Multimodal Large Language Models for Real-Time Knowledge-Grounded Question Answering

Dr. K. Sujatha

The exponential growth of heterogeneous digital information across structured and unstructured repositories presents a critical challenge for large language models (LLMs): the inability to access and reason over dynamically evolving knowledge without costly model retraining. This…

View free PDFSource page