Skip to content

Science

How Tejesha estimates personality from language

The words people choose carry measurable signals about their personality. Those signals are real but imperfect, so Tejesha reports how confident it is in every estimate.

The Big Five

Decades of research on how people describe each other converged on five broad traits: openness, conscientiousness, extraversion, agreeableness and emotional stability, sometimes described by its opposite, neuroticism (Goldberg, 1990; John and Srivastava, 1999).

Each trait is a continuum rather than a category. Most people sit near the middle on most traits, and neither end is better than the other. A person lower in extraversion is not a weaker communicator; they often simply prefer a different kind of conversation.

What higher and lower levels of each trait tend to look like
TraitHigherLower
OpennessCurious and imaginative, drawn to new ideasPractical and grounded, prefers what is proven
ConscientiousnessOrganized and thorough, plans aheadFlexible and spontaneous, comfortable improvising
ExtraversionOutgoing and energized by peopleReserved and reflective, prefers smaller groups
AgreeablenessWarm, cooperative and considerateDirect, skeptical and focused on the task
Emotional stabilityCalm under pressure and slow to worryFeels stress keenly and stays alert to risk

Why traits rather than types

Many workplace tools sort people into a small number of types. Type labels are easy to remember, but they force people who differ by degree into separate boxes, so two people with nearly identical answers can land in different categories.

Researchers have raised this concern about popular type indicators for years, including evidence that many people receive a different type when they take the same test again (Pittenger, 2005). Trait measures avoid the problem by reporting where someone falls on each dimension, and they are the standard in personality research.

What research says about personality and language

How people write is fairly stable over time and is linked to their personality (Pennebaker and King, 1999). Large studies have mapped which words go with which traits (Yarkoni, 2010), and automated models can estimate traits from conversation, essays and social media posts (Mairesse et al., 2007; Park et al., 2015). In one large study, predictions from people's Facebook Likes were more accurate than personality judgments made by their friends (Youyou et al., 2015).

The signal is real, and it is modest. A meta-analysis of studies predicting the Big Five from social media found correlations with people's own questionnaire scores of roughly 0.3 to 0.4, depending on the trait (Azucar et al., 2018). Estimates at that level track personality clearly across many people, yet can be well off for any single person. That gap is why Tejesha shows a confidence score for every trait instead of presenting estimates as fact.

Method

The signals Tejesha looks at

No single signal is reliable on its own. Tejesha combines four kinds and lowers its confidence when they disagree.

Word choice

Categories of words that research links to personality, such as social and emotional language, and markers of certainty or hedging.

Research: Pennebaker and King, 1999; Yarkoni, 2010

Structure and style

Sentence length and complexity, formatting habits, and how a person organizes an argument.

Research: Mairesse et al., 2007

Consistency over time

Writing compared across months, to tell lasting tendencies apart from a single mood or occasion.

Volume and agreement

How much text there is and whether the signals point the same way. Little text or conflicting signals lower the confidence score.

How to judge accuracy claims, including ours

Personality prediction is easy to oversell. Before trusting any accuracy figure, ask these four questions.

What was it compared against?
Look for validated personality questionnaires completed by the same people whose writing was analyzed.
How many people, and who were they?
Small or unusual samples can make results look better than they will be for your contacts.
Was it tested on new people?
Results should come from people the model was not built on. Otherwise they measure memory, not prediction.
Is uncertainty reported?
An average accuracy figure hides how far off an estimate can be for one individual.
Ask us about our methodology

We are glad to walk through how Tejesha is evaluated, limits included.

References

  1. Azucar, D., Marengo, D., & Settanni, M. (2018). Predicting the Big 5 personality traits from digital footprints on social media: A meta-analysis. Personality and Individual Differences, 124, 150-159.
  2. Donnellan, M. B., Oswald, F. L., Baird, B. M., & Lucas, R. E. (2006). The Mini-IPIP scales: Tiny-yet-effective measures of the Big Five factors of personality. Psychological Assessment, 18(2), 192-203.
  3. Goldberg, L. R. (1990). An alternative "description of personality": The Big-Five factor structure. Journal of Personality and Social Psychology, 59(6), 1216-1229.
  4. John, O. P., & Srivastava, S. (1999). The Big Five trait taxonomy: History, measurement, and theoretical perspectives. In L. A. Pervin & O. P. John (Eds.), Handbook of personality: Theory and research (2nd ed., pp. 102-138). Guilford Press.
  5. Mairesse, F., Walker, M. A., Mehl, M. R., & Moore, R. K. (2007). Using linguistic cues for the automatic recognition of personality in conversation and text. Journal of Artificial Intelligence Research, 30, 457-500.
  6. Park, G., Schwartz, H. A., Eichstaedt, J. C., Kern, M. L., Kosinski, M., Stillwell, D. J., Ungar, L. H., & Seligman, M. E. P. (2015). Automatic personality assessment through social media language. Journal of Personality and Social Psychology, 108(6), 934-952.
  7. Pennebaker, J. W., & King, L. A. (1999). Linguistic styles: Language use as an individual difference. Journal of Personality and Social Psychology, 77(6), 1296-1312.
  8. Pittenger, D. J. (2005). Cautionary comments regarding the Myers-Briggs Type Indicator. Consulting Psychology Journal: Practice and Research, 57(3), 210-221.
  9. Yarkoni, T. (2010). Personality in 100,000 words: A large-scale analysis of personality and word use among bloggers. Journal of Research in Personality, 44(3), 363-373.
  10. Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036-1040.

Try the Big Five on yourself

Take the free 20-question test built on the public domain Mini-IPIP inventory. It takes about three minutes.