Papers
Topics
Authors
Recent
Search
2000 character limit reached

Building the Ipseome: Large, Free, Open, Human Identity Data

Published 2 Jul 2026 in cs.DL | (2607.02488v1)

Abstract: Shared data accelerates scientific progress. Here, I describe the ipseome -- the largest free and open dataset on the topic of human identity. The dataset is designed as reusable research infrastructure, with publicly accessible data repositories, documented measurement procedures, and versioned files for cumulative research on identity. First, I present the motivation for and the ipseological principles driving construction of the ipseome. Then, each component is introduced and discussed. Finally, I summarize the current state of progress toward the ultimate goal.

Authors (1)

Summary

  • The paper introduces the ipseome as a comprehensive resource that quantifies human identity using vast self-authored textual data.
  • It employs robust methodologies including daily surveys, longitudinal tests, and global social media analysis to capture identity dynamics.
  • The approach underpins reproducible AI and social science research by enabling cross-cultural, temporal, and large-scale studies of selfhood.

Building the Ipseome: Large, Free, Open, Human Identity Data

Introduction

The paper "Building the Ipseome: Large, Free, Open, Human Identity Data" (2607.02488) presents a comprehensive, systematically constructed resource—the ipseome—that seeks to catalog and quantify human identity at scale, using personally expressed identity (PEI) text. By compiling diverse, high-volume, openly accessible datasets, this work establishes an empirical foundation for the emerging field of ipseology, which is defined as the quantitative and computational study of selfhood and identity elements via self-authored language.

Motivation and Principles of the Ipseome

The author identifies critical deficits in social science infrastructure related to identity research: absence of well-curated, large-scale, longitudinal, and cross-cultural datasets reflecting self-ascribed identity. Analogous to the genomics and connectomics infrastructure in biology and neuroscience, the ipseome is advanced as a catalog of human selves, operationalized as linguistic data. The paper enumerates key driving principles for the ipseological approach:

  • Language as data: All elements of the ipseome are rooted in explicit self-descriptive language, capturing both content and the choices of omission.
  • Self-authored self-descriptions: Prioritization of naturalistic, participant-driven text rather than externally categorized responses.
  • Temporal centrality: Persistent, continuous data collection captures identity dynamics.
  • Scale: Emphasis on large samples, wide demographic and geographic axes, and repeated measurement for precision and generalizability.
  • Empirical foundation first: Quantitative description is prioritized over theoretical presuppositions, enabling future theory to be constrained by robust data.

Description of Major Ipseome Components

Ipseity Daily

Ipseity Daily is an automated, persistent panel survey that samples U.S. adults daily, querying binary (Yes/No/Skip) responses to randomly presented identity signifiers (words, phrases, or emojis). At writing, it includes 707 signifiers and yields 1,680 observations per day (80 per respondent × 21 respondents). The survey is funded for at least one million observations and is openly available. This infrastructure enables granular tracking of the prevalence, co-occurrence, and temporal evolution of identity signifiers across the U.S. population.

JJJ Pro Who am I?

This dataset collects free-text self-descriptions modeled on the Twenty Statements Test (TST) from a representative national sample, including a longitudinal component via repeated respondents. All responses are made publicly available under a CC BY-NC-SA license, including demographic variables and full verbatim text (with minimal redaction for names only). The dataset covers broad age, gender, and ethnicity strata, with open microdata access for external validation and reuse.

HINENI: Human Identities Across Nations of the Earth Ngram Investigator

HINENI draws from Twitter profile biography data collected globally over a twelve-year interval. Employing random sampling and geolocation mapping, the dataset enables cross-national and longitudinal prevalence tracking of identity signifiers. HINENI is the first infrastructure to support annual census-level quantification of identity terms in social self-presentation across thirty-two nations. The data support fine-grained analysis of identity trends, such as political self-ascription's increasing incidence worldwide.

As a predecessor to HINENI, this dataset isolates U.S.-based Twitter users and tracks annual changes in signifier prevalence from 2015–2023, using comparable collection and aggregation methodologies. An online dashboard provides raw and visualization access.

Words You Today

This application supports individual longitudinal tracking of binary responses to identity signifiers, yielding personalized, anonymized timeseries data. The design allows for the study of intra-individual stability and change in identity signifiers at a fine-grained temporal resolution.

Methodological Rigor and Data Governance

All datasets are collected with explicit informed consent, robust anonymization, and careful manual review to eliminate direct identifiers (especially names). Notably, the policy for chatbot-generated or "invalid" responses is inclusivity, maintaining maximal transparency and provenance for downstream filtering and analysis.

Payment and recruitment protocols (e.g., Prolific platform) are specified in detail, with adequate compensation and demographic representativeness reported. The collection processes are positioned as modular and replicable, and the infrastructure is explicitly designed for cumulative research and open science.

Analytical and Research Implications

The ipseome enables a wide array of research questions in computational social science, social psychology, digital sociology, and AI:

  • Population-level prevalence and trend analysis of identity categories (e.g., sexuality, politics, occupation) using scalable, high-frequency measurement [handzlik_hineni_2024].
  • Individual- and cohort-level identity change and survivability, leveraging repeated measures and longitudinal sampling [vahabli_identity_2025].
  • Cross-cultural and cross-national comparisons of identity self-presentation, moving beyond small convenience samples (cf. [rhee_spontaneous_1995]) to hundreds of millions of bios [handzlik_hineni_2024].
  • Identity clustering and co-expression within social networks and within individuals across time, enabling rigorous study of salience, hierarchy, and social influence [tucker_pronoun_2023; maier_online_2026].
  • Benchmarking and development of natural language processing models uniquely suited for social identity recognition, classification, and evolution analysis.

The approach also offers an empirical platform for testing and revising classic and contemporary theories of identity structure and formation (cf. [stryker_symbolic_1980]), resituating the study of self within a thoroughly data-rich, computational paradigm.

Numerical Results and Claims

The author reports the largest openly available datasets of self-authored identity descriptions:

  • Ipseity Daily generates approximately 1,680 observations per day with planned continuation for at least one million total observations.
  • HINENI spans twelve years and thirty-two nations, covering hundreds of millions of Twitter profile bios.
  • JJJ Pro Who am I? provides annual, longitudinal, full-text TST-like self-descriptions in a representative U.S. sample.

Significantly, using this data, the author claims that the prevalence of political identity self-ascription has increased in 30 out of 32 countries with adequate data [handzlik_hineni_2024; clemente_online_2024]. This cross-national politicization of identity is documented as a robust, quantifiable trend.

The data architecture enables quantification not only of category prevalence but also of signifier survivability, co-occurrence, and longitudinal profile dynamics.

Theoretical and Practical Implications for AI

From a theoretical perspective, the ipseome enables the formulation and validation of identity models amenable to computational inference. For AI research, the resource supports the construction of large-scale, up-to-date training and evaluation corpora for social identity tasks, ranging from self-concept labeling, stance detection, user profiling, and temporal modeling. The open, versioned, and cumulative nature of the datasets facilitates reproducibility, benchmarking, and the development of robust models for automated identity recognition—a key challenge in both sociotechnical and responsible AI research.

Long-term, the ipseome framework provides a model for quantitatively studying high-dimensional, dynamic, socially embedded human characteristics, suggesting a pathway for bridging social science and AI at infrastructure and methodological levels. The data can inform adaptive, context-sensitive AI systems attuned to dynamic user identities, mitigate bias, and enable new forms of human-machine interaction that are sensitive to the evolution of self-presentation and group belonging.

Future Directions

The infrastructure supports continual expansion in terms of:

  • Broader demographic, geographic, and linguistic inclusion.
  • Integration with other behavioral and contextual digital traces.
  • Extension to non-U.S. populations and harmonization with global survey efforts.
  • Iterative improvement of signifier lexica and data collection processes for emerging identity categories.
  • Development of cross-disciplinary collaborations linking computational social science, behavioral science, and AI.

The modular principles and open data philosophy position the ipseome as a replicable model for other domains wherein fine-grained, longitudinal, open data can transform theory and empirical research.

Conclusion

"Building the Ipseome" (2607.02488) represents a significant advance in the empirical study of human identity by developing a scalable, openly available infrastructure for ipseology. Through persistent, large-scale, language-centered data collection, the ipseome operationalizes human selfhood as an observable, quantifiable, and dynamic phenomenon. The resource recalibrates the evidentiary standards for social identity research, supports a broad spectrum of AI and analytic applications, and provides a generalizable model for open, cumulative research in social science and AI domains.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.