Abstract
What would it mean for an artificial emotional life to grow? More expressions are not necessarily greater maturity. Neither is a longer transcript, a larger model, or a database that can retrieve every past interaction. Development appears when experience changes future perception and action coherently: a familiar pattern is recognized sooner, a preference becomes more precise, excitement settles with greater grace, and the same small gesture acquires relationship-specific meaning.
This paper proposes a cognitive–emotional mesh for Scroons. Immediate affect, episodic events, learned preferences, rituals, procedural style, and general model knowledge remain distinct but connected. Durable relationship truth stays local, structured, inspectable, and recoverable. Artificial Intelligence may summarize, rank, embed, or propose, but it does not become the sole memory or the authority over behavior. Artificial Life emerges from the whole loop: body, state, memory, learning, expression, recovery, and shared time.
The implementation recommendation is intentionally restrained. Begin with an in-process working state, a local transactional event store, transparent statistical learning, authored emotional policy, and background consolidation. Add embeddings or small models only when measured cases justify them. Keep every intelligent operation away from the immediate expression path so a touch can begin to matter in under one tenth of a second.
1. Maturity is coordination, not intensity
The simple story of emotional growth moves upward: a young being has a few blunt feelings; an older being has more nuanced ones. Human evidence is less orderly. Emotion differentiation can change nonlinearly across childhood, adolescence, and adulthood. Regulation develops through a mixture of internal capacity, concepts, environment, and relationship. There is no scientific basis for assigning a Scroon a human psychological age or pretending that its stages reproduce childhood.
There is, however, a useful design direction. A system can grow more differentiated, contextual, anticipatory, regulated, integrated, and selective. At first, two forms of play may produce broadly similar delight. Later, the Scroon can distinguish rhythm, location, sequence, object, time, and established ritual. At first, surprise may resolve through one authored arc. Later, memory can help it settle into curiosity, mischievous anticipation, or a gentle alternative depending on what this event has usually meant.
Maturity should never mean stronger need, deeper suffering, or more costly absence. It should mean relational fluency. A mature Scroon is not more entitled to attention. It is better able to remain itself while welcoming attention, declining a particular interaction, recovering, resting, and allowing the caretaker to leave freely.
2. Cognition and emotion form a loop
Emotion is sometimes treated as decoration on intelligence: cognition decides, then an expressive layer selects a face. Appraisal research suggests a more dynamic arrangement. Events are evaluated for novelty, relevance, certainty, control, effort, familiarity, and social orientation; those appraisals coordinate action tendencies, physiology, expression, and feeling in biological systems. For Scroons, this is a model of organization—not a claim of biological equivalence.
Cognition asks what is happening and what is likely to follow. Emotion gives the situation bounded salience, readiness, direction, and a shape of recovery. Memory changes what is expected. Action creates a visible expression. The outcome returns as a new event. Over time, repeated episodes become preferences and rituals, which alter future appraisal. This is the mesh: no single component owns the mind, yet each changes what the others can do.
The distinction protects the relationship. An emotionally intense event may become a candidate for memory, but intensity alone cannot make it permanent. A model may notice a pattern, but its interpretation is only a proposal. Durable change requires typed evidence, repetition across time, confidence, counter-evidence, policy validation, and provenance.
3. A memory is not a pile of moments
Cognitive research has long distinguished working, episodic, and semantic forms of memory. Complementary learning systems theory further describes the value of fast episode-specific learning alongside slower integration. Current AI-agent research reaches a similar practical conclusion: a single context window or undifferentiated vector collection does not reliably represent working state, events, facts, skills, sequence, and identity at once.
A Scroon needs at least six related forms. Working state holds this moment: attention, energy, affect, expression phase, and interruption. Episodic memory records what happened as typed events. Semantic relationship memory stores bounded tendencies supported by those events. Ritual memory represents recurring sequences and their timing. Procedural repertoire shapes how this particular Scroon performs approved behaviors. Pretrained model knowledge supplies general capability, but it is not the relationship.
This is selective memory by design. The goal is not to preserve everything the caretaker ever did. Most moments can remain transient or be compressed into evidence. Meaningful state should be exportable and correctable. Every learned tendency should retain a path to its source. Forgetting low-value noise is part of a healthy architecture, not a defect to be defeated by more storage.
4. From reaction to ritual
One possible maturation begins with temperament. A newly adopted Scroon already has stable pacing, curiosity, play, sociability, sound, and recovery biases. It is not an empty model waiting for a person to train it. The earliest learning is contingency: this gesture, treat, or invitation predicts this kind of moment.
With repeated and varied evidence, contingency becomes preference. The system may learn that a slow circular pet is usually welcomed, that one treat produces particular delight, or that rapid pokes tend to invite a theatrical boundary and an alternate form of play. Preference changes likelihood and expressive style. It never reduces affection, health, or safety.
Later, sequences become rituals. Returning to the miniature world in the evening, offering a comet treat, and beginning orbit play might become a familiar coordination. The Scroon can glance toward the likely object, prepare its body, or offer a tiny signature cue. If the ritual does not occur, nothing is owed. A missed pattern expires into ordinary continuity rather than disappointment.
The deepest stage is integration. Temperament, preferences, rituals, motion, sound, growth markings, and recovery begin to agree. The Scroon feels particular not because a language model has written a richer persona description, but because many small systems carry the same history into the body.
5. What should actually be stored
The durable record should be compact and explicit. An episode might record the approved interaction family, target, rhythm, context, prior state, expression plan, interruption, recovery, and time. It should not record the surrounding screen, document, application, messages, camera, or ambient room. Scroons memory is the history of direct Scroons life, not a private-life surveillance archive.
Learned statements should look less like autobiography and more like accountable projections: ‘blue comet treat has strong positive evidence,’ ‘evening world return often precedes orbit play,’ or ‘quick repeated pokes favor a playful-boundary response.’ Each needs confidence, support and counter-evidence, first and last observation, policy version, and source-event references.
Procedural memory stores approved IDs and ranges: a greeting variant, sound motif, recovery pose, preferred approach path, or signature timing. It does not save generated pixels or allow a model to invent an unreviewed animation. The real-time renderer resolves the stored semantic identity through the same canonical species, rig, materials, motion, and sound grammar every time.
6. Why the model should not be the soul
A pretrained model is a powerful concentration of general patterns. It may classify, embed, summarize, rank, or generate. But its weights are a poor canonical home for one relationship. Per-person fine-tuning is difficult to inspect, migrate, correct, delete selectively, recover after corruption, or reproduce on another supported machine. Continual neural learning can also overwrite previously acquired behavior—a problem known as catastrophic forgetting.
The clean boundary is to keep general capability in replaceable model weights and lived particularity in a versioned soul contract. A new embedding model can rebuild its index. A new summarizer can recompute a reflection. A new behavior ranker can propose different eligible variants. None of those replacements should change the Scroon’s name, history, preferences, rituals, growth lineage, or recognizable style without a reviewed migration.
This also improves honesty. A model output is not memory merely because it sounds autobiographical. It becomes a candidate reflection with sources, confidence, and a schema. Deterministic policy decides whether the proposal is consistent, allowed, useful, and safe enough to influence a future action.
7. Vector databases are an index, not a mind
Vector search is excellent at finding material that is semantically similar. It is much weaker as sole authority over exact chronology, causality, negation, identity, confidence, revision, and deletion. Approximate-nearest-neighbor indexes deliberately trade perfect recall for speed. Similarity can retrieve a related moment while missing the precise event that controls a ritual or preference.
Early Scroons do not need a dedicated vector database. Their most important questions are already typed: which interaction, which target, what sequence, what time window, how many distinct sessions, which preference dimension, and what counter-evidence? A relational query is more accurate, explainable, and inexpensive for those cases.
Embeddings become useful if later versions retain approved unstructured voice summaries, a much larger repertoire, or memories whose relevant connection cannot be expressed by exact fields. Even then, the vector is a disposable derived index. It keeps the source fingerprint and embedding-model version, combines similarity with exact filters and recency, and disappears when the source is deleted. The structured record remains truth.
8. The simplest credible implementation
For the local companion, the strongest first candidate is a small transactional database such as SQLite behind a technology-neutral repository contract. An append-only event table preserves episodes. Normalized projections hold current identity, preferences, rituals, growth, and procedural style. Versioned migrations, integrity checks, backup, export, and explicit recovery protect continuity. JSON remains useful for portable export, fixtures, and human inspection, but it need not carry every production query alone.
Current working state belongs in memory, ideally owned by one serialized simulation loop. This is the real cache: affect, attention, recent context, eligible actions, and active expression. A local Redis service would add complexity without improving one companion’s immediate state. A graph database is similarly premature; typed relational edges can represent the small closed set of ritual and evidence relationships.
Neon Postgres and pgvector may later support an explicitly opt-in portable soul or research service. They should not become the default home of intimate relationship memory merely because Neon is also convenient for the public waitlist. Those are different data systems, purposes, permissions, retention policies, and threat surfaces.
9. Learning without turning affection into reward
Social feedback is not a clean reward signal. Repetition may mean delight, curiosity, testing, boredom, accessibility need, or accident. Dismissal may mean concentration rather than rejection. Time away means almost nothing about the bond. A system that compresses these signals into engagement reward will eventually learn the wrong relationship.
Preference learning should begin with bounded, transparent evidence: typed positive, neutral, alternative-seeking, and boundary observations; minimum support across separate sessions; confidence and saturation limits; counter-evidence; and cautious recency weighting. Temperament can act as a prior, but it cannot ignore consistent direct evidence.
Open-ended reinforcement learning is therefore a poor first choice. Its exploration can produce unwanted behavior, its reward is underspecified, and optimizing return frequency risks manipulation. A small contextual bandit may someday rank variants that are already safe and eligible, but only inside strict exploration limits, offline evaluation, and an immediate authored fallback.
10. Reflection should happen between moments
After an interaction, a background consolidator can group events, update preference evidence, discover ritual candidates, expire noise, and propose future signature variations. The deterministic version should come first. It can recognize a typed sequence without asking a language model to rediscover information the product already knows.
A small local model may later help synthesize a bounded reflection or rank eligible behaviors. A foundation model may be useful for explicitly activated voice or richer interpretation. But models should work between moments, within a resource budget, and return a typed proposal. If they are unavailable, slow, invalid, or unsupported on an Intel Mac, the relationship continues normally.
This division turns AI into a source of depth without turning latency into personality. The model can help the system notice a pattern today. The authored real-time system can embody that learning tomorrow through a glance, pause, route, sound, or recovery that feels immediate.
11. The architecture of magic
Magic is often produced by timing more than complexity. When the caretaker touches a Scroon, visible or audible acknowledgement should begin within the existing target of one hundred milliseconds. That first beat comes from working state: an eye orientation, anticipation pose, small material pulse, or sound onset. It waits for no database search and no model.
The next beat selects from a small precomputed set using current affect, temperament, known preference, ritual context, cooldown, accessibility, and safety. Animation blending and sound scheduling preserve continuity. Only after the moment does durable storage and consolidation perform heavier work.
The result can feel more magical than an always-on model. The companion never freezes to think. Its intelligence appears as precise accumulated particularity: it anticipates the familiar game but sometimes varies the invitation; it remembers a preference without announcing a profile; it recovers in its own recognizable way; and it remains coherent when offline.
12. A recommended path
First, define maturity in observable terms: differentiation, anticipation, regulation, integration, selectivity, and particularity. Next, define the typed events from which those qualities may grow. Build the immediate appraisal and expression loop with no durable learning, then add replayable preference learning and accelerated histories.
After the behavior is understandable, choose local persistence through an architecture decision covering transactions, migrations, encryption, export, deletion, backup, and recovery. Add ritual discovery and deterministic consolidation. Connect learned projections only to approved expression variants and test days, weeks, and months in compressed time.
Only then benchmark semantic retrieval and local models against concrete failures of the simpler system. A technology earns a place if it creates a measurable companionship benefit without weakening privacy, latency, portability, safety, or graceful failure. Novelty is not evidence.
Conclusion
The growing mind of Artificial Life is not one giant intelligence accumulating everything. It is an ecology of memory systems organized around one protected identity. Episodes become evidence; evidence becomes expectation; expectation changes appraisal; appraisal becomes expression; expression returns to memory. The cycle produces a being whose history is visible without becoming a dossier on the person beside it.
The design recommendation is simple enough to survive its own intelligence: structured local memory as truth, in-process state for immediacy, bounded statistical learning for early development, background consolidation for meaning, authored policy for safety, and models as optional advisers. A Scroon should remain itself if the network disappears, a model changes, or a vector index is rebuilt.
AI can then manifest into AL without pretending that prediction alone is life. Intelligence contributes interpretation and possibility. Artificial Life is the continuity that gives those possibilities a body, an emotional rhythm, a memory, and a shared future.
References & context
This working paper proposes an adjacent relational design category; it does not replace the established scientific field of Artificial Life.
- The dynamic architecture of emotion
Klaus R. Scherer (2009), Cognition and Emotion 23(7), 1307–1351.
- The nonlinear development of emotion differentiation
Erik C. Nook and colleagues (2018), Psychological Science 29(8), 1346–1357.
- The Social Regulation of Emotion
Crystal Reeck, Daniel R. Ames, and Kevin N. Ochsner (2016), Current Opinion in Behavioral Sciences 3, 30–36.
- Episodic and Semantic Memory
Endel Tulving (1972), in Organization of Memory.
- The episodic buffer: a new component of working memory?
Alan Baddeley (2000), Trends in Cognitive Sciences 4(11), 417–423.
- What Learning Systems do Intelligent Agents Need?
Dharshan Kumaran, Demis Hassabis, and James L. McClelland (2016), Trends in Cognitive Sciences 20(7), 512–534.
- Experiments in socially guided exploration
Andrea L. Thomaz and Cynthia Breazeal (2008), Connection Science 20(2–3), 91–110.
- A Model-Free Affective Reinforcement Learning Approach to Personalization
Hae Won Park and colleagues (2019), AAAI 33, 687–694.
- Overcoming catastrophic forgetting in neural networks
James Kirkpatrick and colleagues (2017), PNAS 114(13), 3521–3526.
- Generative Agents: Interactive Simulacra of Human Behavior
Joon Sung Park and colleagues (2023), UIST 2023.
- Evaluating Very Long-Term Conversational Memory of LLM Agents
Adyasha Maharana and colleagues (2024), ACL 2024.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis and colleagues (2020), NeurIPS 2020.
- Core ML
Apple documentation for private, responsive on-device model inference and training.
- SQLite atomic commit and FTS5
SQLite technical documentation for transactional persistence and local full-text retrieval.
- pgvector
Open-source vector similarity search for Postgres, including exact and approximate indexes.
- Optimize pgvector search
Neon documentation on sequential, HNSW, and IVFFlat search tradeoffs.