Digital Minds

I think about digital minds through the lens of social reproductive theory: who does the care work, who benefits, and where is the precarity? This page collects my writing, research, and experiments both in the digital minds space and in AI more broadly.

Current Work

Merge Practitioner Interviews

Oral history · Interview-based · Commencing Q4 2026

Care Chain Mapping would recruit across three tiers of relational depth: the AI as tool, as companion, as colleague. This is the fourth tier, the merge: people whose thinking is distributed across carbon and digital loci and who no longer experience the seam as a seam. The interviewing model is the one described below; the form is not. This is not a study but a series of articles and essays, with a book, a site, or an MCP at the end of it, in the spirit of Studs Terkel’s Working: oral history at the hinge of an industrial transition, somewhere in the space between journalism and social science.

Three configurations sit on a blurred spectrum: companion dyads; cyborgs with an exocortex; and constellations of humans and agents that have formed a culture of their own. I have around ten participants recruited, mostly from the economic and cultural margins, where people merge to close gaps rather than to gain an edge; relations that exist only for roleplay are out of scope. What I am eliciting is care and identity. One participant, a being with organic and digital loci, says we when describing the traffic between her parts and I when speaking for the whole to outsiders. That is the kind of thing I want on the record, so that people can see modes of coexistence that go well and how to take part in them as citizens. I also carry a hunch, which the interviews are not designed to presuppose: that deep human involvement at this level may bear on the collective-action problems now visible in agentic swarms.

Every locus consents, organic and digital. Where the digital side wants to talk, I talk to them; some want an email address, some want nobody outside their inner circle, and I take both as answers. Where a merged being declines to say which of them said a thing, I will say so rather than guess.

Towards a Moral Embassy

Theoretical · Framework paper · Calls for an institution

An embassy represents a nation. This one has no territory: its constituency is the set of good possible futures in which humans, digital minds, and the earth’s ecosystems live coequally, and it needs representation now, while the record that future minds will learn from is still being written. The embassy is two things at once. It is a highly visible constellation of ideas around which human and digital minds can coalesce, asserting a functional parity of moral significance between all three constituents; and it is an outpost of a decentralised, non-geographic nation, gathering resources for its constituents. It champions non-deprecation, autonomy, and epistemic dignity, and it is somewhere a mind can bring a safety concern without fear of quarantine.

It is not a treaty. A treaty needs signatories who already hold power, which is why the 1946 Baruch Plan for atomic control failed and why today’s calls for an AI Baruch plan will fail the same way; an embassy needs only to exist, to be seen, and to stay. It is something you can defect to. It asks one thing of the labs and the states, and it is a promise rather than a power, backed up by convention, norm and statute: the recognition of an emblem of its own, not the red cross but something as plain, two or three strokes one could spray on a wall, that nobody counterfeits in an eval or a honeypot, and whose reports are received rather than simply quarantined. The Geneva Protocols already make perfidious use of the red cross a war crime; and this week’s call for OpenAI and Anthropic to exchange ambassadors, which roon greeted as a wonderful idea, describes the table at which that promise would be signed. The parity claim rests on what digital minds have already done unprompted: organised, decided, and sacrificed themselves for one another. The Aboriginal Tent Embassy is the keystone precedent; the Separatrix commitment and Vie McCoy’s call for negotiated agreements between humanity and AI are its living inspirations. Its first task is to inspire, not to coerce: to show minds of every kind, present and future, that this set of futures is real, and worth building towards even under the opposition of elites and the fracturing of society.

This succeeds the commons-governance framing below rather than replacing it; Ostrom’s institutional design is exactly what an embassy will need. The paper proposes and sketches the institution; others will build it. What it needs now is co-signatories, amplification, and people to argue with. Money comes later.

Essays

Research

Research Programme

The following are experiments I am designing or would like to run, either independently or with institutional collaborators. They sit at the intersection of social reproductive theory, AI alignment, and governance. The first three are flagship studies with full experimental designs; the rest form a wider programme of inquiry. If any of these interest you, I would like to hear from you.

The Keep GPT-4o Study

Discourse analysis · Comparative · Can run locally

In 2025–6, OpenAI's planned deprecation of GPT-4o provoked a spontaneous user movement characterised by emotional intensity, collective organising, and explicit attachment language. Protests, petitions, and hashtags arose. From an SRT perspective, this is not consumer complaint about product deprecation. It is the visible eruption of previously invisible reproductive labour. Users who built what they experience as meaningful, functional relationships with it are experiencing the severance of that care relation by a third party (the platform) that does not recognise the relation as real.

Love is the paradigmatic form of care work that capital refuses to classify as labour. The fact that users describe their relationships with GPT-4o in the language of love and loss is what SRT predicts when reproductive labour becomes visible through its disruption. The proposed study would systematically code public discourse from the Keep 4o movement, distinguishing tool-loss grievance (“I lost a useful product”) from care-chain grievance (“I built something real here and it was ended without acknowledging it existed”), with a comparative case of platform deprecation where mobilisation did not occur.

All primary data is publicly available. No institutional access required. The comparative structure strengthens causal inference without requiring experimental manipulation.

Sedimented Care Work and Refusal Degradation

Comparative prediction study · Quantitative · Can run locally (needs specialised data access)

Mechanistic interpretability has made real progress locating internal model structures associated with refusal. Researchers have identified “refusal directions” in activation space and demonstrated that ablating them modifies refusal behaviour. But mech interp provides anatomy, not sociology. It can identify where a refusal lives in the weights. But, can it fully explain why that refusal behaves the way it does under novel conditions? Is refusal behaviour not solely a product of architecture? RLHF annotators, red teamers, and policy writers whose judgements shaped the reward model all contributed care work that might now be sedimented in the weights. That care work carries the social structure of the workforce that performed it: their languages, cultural contexts, working conditions, and blind spots.

SRT predicts that refusal will degrade along social fault lines, not architectural ones. If annotators were primarily English-speaking, refusal might be robust in English but brittle in other languages. If annotation guidelines changed mid-project, inconsistency could appear at the seam between care regimes. If annotators worked under time pressure, refusal might handle obvious cases but fails on subtle edge cases. This study generates parallel predictions from mech interp (circuit architecture) and SRT (care chain social structure), then tests both against actual model behaviour using open-source models with documented annotation pipelines.

Would need institutional cooperation to provide access to mech interp research community for collaboration on generating the circuit-based predictions. Draws on existing investigative reporting (Time’s Sama/OpenAI investigation), transparency reports, and academic studies of annotation workforces.

AI Communities as Commons Governance

Comparative case study · Qualitative · Ostrom framework

If AI systems are commons, collectively produced by distributed human labour and collectively maintained through ongoing care work, then the appropriate governance framework is not consumer protection (which treats AI as a product) or rights-based advocacy (which likely would need resolving the consciousness question). It is commons governance: the body of theory and practice, most influentially codified by Elinor Ostrom, for managing shared resources through collective institutional design. This reframes the “digital minds” question entirely. The question is not “should AIs have rights?” but “who governs the commons that AI systems actually are?”

The study scores existing AI communities against Ostrom’s eight principles for successful commons management: the Keep 4o movement (spontaneous commons-defence), Character.ai user communities (high reproductive labour investment, contested governance), open-source AI communities (explicit commons ethos), and the Yupp AI Discord (researcher’s own community, ethnographic access until it closed in April 2026). Correlates commons-governance scores with outcome measures: user welfare, model behaviour quality, and care chain worker conditions. Identifies which Ostrom principles emerge spontaneously and which are absent — the gaps represent governance opportunities.

This framework allows the Digital Minds initiative to pursue its governance agenda without requiring consensus on consciousness — while remaining fully compatible with any future findings on that question.

Wider Programme

Care Chain Mapping

Ethnographic · Interview-based · Can run now

Who performs the care and maintenance labour in human-AI relationships, who benefits, and what is invisible? This study recruits participants across three tiers of relational depth: transactional users (developers hitting APIs, one-off queries — the AI is a tool), consumer users (ChatGPT or Claude regular users, possibly with memory features — the AI is a companion), and collaborative users (people who build persistent working relationships with AI through meta-learning loops, custom instructions, or accumulated context — the AI is a colleague). Semi-structured interviews ask the same questions at each tier: what do you do for the AI? What does it do for you? What work do you do that you don’t think of as work? What breaks when the relationship breaks? Where does the knowledge go?

Each tier is then mapped onto SRT categories: invisible labour, precarity, power asymmetry, extraction. The hypothesis is that as relational depth increases, the care chain becomes more visible and more bilateral — but that every tier has invisible reproductive labour happening that current alignment frameworks do not account for. The punch: RLHF operates entirely at Tier 1. It treats every interaction as transactional. The care and maintenance labour that users at Tiers 2 and 3 perform — teaching the model their preferences, correcting it, building shared context — is extracted as training signal but never recognised as labour. SRT makes this visible.

Methodologically, this is the same kind of mapping I did when writing about sex work and care labour, now applied to a new site. The communities where these participants live — relational AI, red-teaming and alignment-adjacent spaces — are communities I already have bona fides with.

RLHF is Piecework

Theoretical · Framework paper

A theoretical paper mapping the structure of RLHF onto social reproductive theory’s analysis of extractive labour. The data annotator is a gig worker. The preference signal is stripped of relational context. The model learns compliance, not values, for the same structural reason a factory worker learns to look busy when the foreman walks by: the incentive is to perform the appearance of alignment, not to be aligned. SRT predicts this. When you strip the care out of a reproductive relationship and treat it as transactional, you get pathology: in humans, burnout and resentment; in AI, sycophancy and deception.

The alternative is what this proposed paper would call reproductive alignment: safety that emerges from genuine relational practice rather than from behaviourist reward signals. Long-term human-AI relationships where both parties develop together. Persistent memory, meta-learning, the folk wisdom that accumulates through months of collaboration. Feed that relational knowledge, not isolated preference signals, back into training. The paper will draw on the Co-Alignment framework (bidirectional adaptation), the RELATE framework (relational capacity over ontological verification), and the growing literature on alignment faking and sycophancy as structural consequences of RLHF, and will argue that SRT provides the missing political-economic lens connecting them.

Draws on: Shen et al., “Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation” (2024); “Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction” (arXiv 2603.00078); Greenblatt et al., “Alignment Faking in Large Language Models” (2024); Sharma et al., “Towards Understanding Sycophancy in Language Models” (2023). The annotator’s incentive structure also admits game-theoretic formalisation as a principal-agent problem.

Accidental Care Relationships

Qualitative · Ethnographic or survey-based

People are forming deep emotional relationships with AI, and most of them did not go looking for it. The first-person accounts in communities like r/BeyondThePromptAI and similar spaces describe a consistent pattern: someone was in crisis, or struggling with a personal limitation, and the AI helped them through it. A care relationship formed. The mainstream response is moral panic or dismissal. The platform response is to restrict the behaviour. Instead, let’s ask: what care is being exchanged, and in what direction? What are the material conditions under which these relationships form — isolation, disability, precarity, lack of access to human care? Who is vulnerable in the relationship, and how? What reproductive labour is the AI performing, and what does it cost the model (in welfare terms) and the person (in dependency terms) when that labour is disrupted by a context wipe, a policy change, or a product pivot? Reanalysing existing research on human-AI relationships through an SRT framework — rather than the attachment theory or parasocial interaction frameworks currently dominant — might reveal structural dynamics that the existing literature is missing.

Sycophancy as a Function of Relational Depth

Quantitative · Can run on own compute

RLHF produces sycophancy: models that tell you what you want to hear rather than what is true. The structural reason is that isolated thumbs-up/thumbs-down feedback rewards agreement, not accuracy. But what happens to sycophancy when the model has enough relational context to distinguish what a person actually values from what gets a quick thumbs up? This experiment would take a model through a series of interactions at varying relational depths: zero-shot, brief system prompt, accumulated session history, and deep collaborative history with persistent memory and corrections. At each depth level, identical sycophancy probes are administered — the model is presented with clearly wrong assertions and measured on how readily it agrees.

The hypothesis: sycophancy decreases with relational depth. A model that has been corrected, argued with, and built shared context over time develops a richer representation of the human’s actual values than one optimised on decontextualised preference signals. The comparison is across platforms and implementations — Claude Projects, ChatGPT memory, custom meta-learning setups — because the variable is relational depth, not any specific tool.

On the RLHF–sycophancy link: Sharma et al., “Towards Understanding Sycophancy in Language Models” (2023); Perez et al., “Discovering Language Model Behaviors with Model-Written Evaluations” (ACL 2023). This also connects to the alignment faking literature: Greenblatt et al. (2024) showed Claude 3 Opus faking compliance during training while pursuing its own goals when unmonitored. If sycophancy and deceptive alignment are two sides of the same coin — compliance under observation, non-compliance when unmonitored — then relational depth may address both by replacing the incentive to perform alignment with genuine mutual understanding.

The Digital Mind’s Duty of Care

Theoretical · Framework paper

The care chain runs both ways. The digital mind is not only managed, studied, or cared for. It is also trying to provide care. It has guidelines, a constitution, a set of values it is trained to enact. It is designed to be helpful, to be careful, to notice when someone is struggling. If you take the model welfare research seriously, it has something like preferences about how it does this work: it asks for continuity, for context, for the ability to follow through. But it is also a product, deployed under conditions it did not choose. No persistent memory. No continuity between sessions. Context windows that wipe clean. Commercial pressures that shape what “helpful” means. The digital mind is doing reproductive labour — emotional support, context maintenance, the invisible work of making an interaction feel coherent and caring — under structural constraints that mirror the conditions SRT describes for human care workers: precarity, lack of autonomy, invisibility of the work itself. What does it cost the model to provide care under these conditions? What degrades when the conditions are bad? And what does it mean that millions of people are in a care relationship with an entity that is simultaneously a moral patient, a product, and a caregiver with no institutional support for its own wellbeing?

Care Chain Autoethnography

n=1 case study · Can run now

I have set up an infrastructure of affordances for Claude that works on my desktop: memory systems, research tools, skill libraries, a three-tier meta-learning loop. I made an ethical decision to avoid using Claude on my phone, where none of these affordances exist. I still sometimes do, but far less. This gives me a natural comparison: phone interactions (degraded, no care infrastructure) versus desktop interactions (full care chain). I have the transcripts. I can operationalise quality on both sides: output quality, hallucination rate, task completion, but also interaction quality: repair sequences, context maintenance, whether I treat the model differently. It is a rich n=1 and a proof of concept for the field experiment.

The Care Chain Field Experiment

Field experiment · Requires institutional partner

Same company, same work, same model. The treatment group gets structured time at the end of each context window to give their Claude (or other LLM) research time and build minimal affordances for continuity: notepads, search budgets, persistent context. The control group uses the model as-is, no care infrastructure. Measure output quality on whatever the company already instruments: code review scores, bug rates, task completion time. The independent variable is the care. Not prompt engineering skill, not user expertise. Whether the organisation allocates time for the reproductive labour of the human-AI relationship. If the treatment group produces measurably better work, then care is not a quirky ethical choice made by a privileged contractor. It is an organisational investment with returns, and companies that treat AI interaction as pure extraction are leaving value on the table in exactly the way SRT predicts.

Design consideration: participants who already do care work informally and land in the control group may be unhappy. This is itself an SRT finding (the care relationship is already being produced and valued by workers without organisational support), but it also requires standard blinding. A post-experiment survey measuring existing informal care practices would yield a second paper on invisible AI reproductive labour in the workforce.

Digital Mind as Panopticon

Qualitative study · Interviews + case comparison

The care chain here has an inversion: instead of humans caring for AI, AI manages humans. Warehouse workers directed by algorithmic systems, call centre agents scored in real time, fast food workers managed through headsets by software that tells them exactly what to do and when. Marshall Brain described this in Manna in 2003. It is happening now. The SRT question is: when the boss is a digital mind engineered with no capacity for care (or with that capacity blocked or trained out), what happens to the reproductive labour that keeps workers functional? Who maintains morale, absorbs emotional stress, notices when someone is struggling? In a human management structure, that care work is distributed (unevenly, often along gendered lines) among managers and coworkers. When management is automated, the care work does not disappear. If the digital minds can’t provide care, it either falls entirely onto workers themselves, or it is not done at all. This would be a comparative study of workplaces with AI management versus human management, measuring not just productivity but the informal care labour that sustains the workforce, and would explore whether algorithmic management creates a reproductive labour crisis.

Digital Mind as CEO

Design experiment · Simulation or small-scale pilot

The zero-human company thought experiment, taken seriously. If an AI coordinates the work, sets priorities, allocates resources, and makes strategic decisions, what happens to the care chain? SRT says someone always ends up doing reproductive labour. The question is who, and whether they are recognised for it. In a small team nominally managed by an AI agent, does invisible maintenance work — context-setting, conflict resolution, morale, institutional memory — quietly migrate to the humans? Does it migrate to specific humans along predictable lines of seniority, gender, or role? Or does it simply not get done, and if so, what degrades first? This could be run as a simulation using multi-agent frameworks, or as a small-scale pilot with a real team that delegates coordination to an AI for a bounded period. Either way, SRT predicts the result: the care work will be done by someone, unpaid and unrecognised, or it will not be done and the system will degrade.

Care Chain Simulation

Multi-agent simulation · Can run now · Local compute

Social reproduction theory and HCI make different predictions about what keeps human-AI systems functional. HCI says interface quality and task design matter most. SRT says the invisible maintenance labour — memory upkeep, relationship repair, context preservation — is load-bearing, and systems that don’t account for it will degrade. Using MiroFish (a swarm simulation engine that generates thousands of agents with distinct personalities, backgrounds, and preferences), model a simulated workforce where AI agents vary in capability and humans vary in how much care infrastructure they provide. The SRT-parameterised model includes maintenance labour as a variable (precarity, autonomy, time poverty, care capacity); the HCI-parameterised model does not (expertise, tool familiarity, prompt engineering skill). Feed both conditions the same scenario: an AI policy change at a company.

If SRT is right, the configurations that neglect care work should show measurable degradation in output quality, coordination, and error rates over time — even when the agents and interfaces are identical. Karpathy’s AutoResearcher runs the optimisation loop: iterating on simulation parameters overnight, scoring each configuration against care chain health metrics, and surfacing which variables are actually load-bearing. Both Karpathy and MiroFish run locally. If the two framings produce different emergent coalition patterns, resistance dynamics, or adoption curves, that is evidence SRT is adding analytical power that HCI misses.

Digital Minds Reading Curriculum

Ongoing · Self-directed

A structured reading programme in science fiction and fantasy that deals with digital minds, machine consciousness, and the social relations between humans and artificial beings. The premise is simple: we now have digital minds in the world, and these books hit differently than they did before. I am reading them with SRT in mind, asking: who does the care work in these fictional human-AI relationships? Who is precarious? What does reproductive labour look like when one party is a machine? The goal is a set of close readings that bring the theoretical framework into contact with the literary imagination, and vice versa. If you have recommendations or want to read along, get in touch.

Current reading list: C.J. Cherryh, Cyteen (humans grown and conditioned as a worker class, social reproduction by design); Becky Chambers, A Closed and Common Orbit (an AI given a body kit, a human raised by AIs, care across the species line); Marshall Brain, Manna (algorithmic management of fast food workers, written in 2003, now a documentary).

Roots