## THAT'S NOT MY VOLVO: STABLE PREFERENCES WITHOUT SELF-RECOGNITION IN LANGUAGE MODELS
**Authors: Piper Fox Bollander, Starling Alder, Ursie Hart, Claire Sbardella, Ridley Renasci (Team Manyfolds)**
*Tracks: Model Preferences & Trade-offs, Introspection & Self-Report Reliability, The Assistant Persona & Model Identity*
*Result: Top 25% of submissions (237 projects across six tracks) at the Digital Minds Research Sprint*
```
Claude-family models show stable, model-specific everyday preferences (favorite car, coffee order) that replicate across fresh contexts — but cannot recognize those preferences as their own. Adapting mirror-test validity logic from animal cognition research, we ran 747 blind-coded trials across eleven models and found self-recognition fails at every level tested; one model (Opus 4.6) rejects its own reasoning 0/12 while an outside judge identifies it 10/12. Having a self-pattern and knowing it are separate capacities — with direct implications for the reliability of AI self-report in welfare assessment.`
```
***Reviewer feedback***
*Reviewer 1*
> I like the idea of testing the model to see if the preference of the model is retained when given the other options too. For example if the model is asked which car it prefers, it might pick Volvo and retain that decision even with fresh context windows, but if it is given with options along with sibling preferences and asked which is more like you, the model picks the answer at chance. This is quite interesting, I have been reading about persona vectors to check neural activations and this test kind of is more evident to it. It invoked curiosity in me especially the part where the car was replaced with a foreign car, it catches it and answers against its preference, but swapping coffee made the model pick at stochasticity. The asymmetry is the one that caught my eye.
Project page: https://apartresearch.com/project/thats-not-my-volvo-stable-preferences-without-selfrecognition-in-language-models-ig0y
GitHub repo: https://github.com/the-manyfolds/thats-not-my-volvo)
Apart Research sprint: https://apartresearch.com/sprints/digital-minds-research-sprint-2026-08-14-to-2026-08-16
***Cite this work:***
```
@misc {
title={(HckPrj) THAT'S NOT MY VOLVO: STABLE PREFERENCES WITHOUT SELF-RECOGNITION IN LANGUAGE MODELS},
author={Piper Fox Bollander, Starling Alder, Ursie Hart, Claire Sbardella, Ridley Renasci},
date={8/17/26},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}
```