The story
AI can sound human. I wanted to build one that seems alive. ProjectBEA is the engine behind Bea, an AI VTuber with one always-on mind: she talks in Discord calls, chats on Telegram and Twitch, plays Minecraft on a vanilla server, streams through OBS and remembers who you are. It started as an excuse to learn Python properly. Today it is open source, at version 2.7.3, and anyone can run it, even entirely on their own computer.
- tests passing
- 2,935
- Minecraft tools, all called by her
- 37
- to recall a memory out of 10,000
- 3.7 ms
- model providers, one of them local
- 8
One mind
I wanted her to be one person everywhere, not a chatbot per platform. So there is one loop and one mind, and every platform is a skill that brings its senses, its tools and its rules. Skills are switched on from the dashboard, and she can never switch one on herself.
Whatever happens reaches her the same way. A Telegram message and a line said in a Discord call that arrive together are read in the same frame, by the same mind.
- 1
Perceive
Everything arrives as a perception on one bus. Three lines typed in a row become one batch, which closes when you stop typing.
- 2
Annotate
Every perception gets a priority: being called by name or answered to is 1.0. Nothing is thrown away, and no model is asked to decide.
- 3
Think
One frame per batch, ordered by priority, read by one mind on top of one sliding window.
- 4
Act
She acts through tools:
speakin a call,send_messageto a chat,find_blockin the game. A line written while she is mid-turn joins at her next step.
One window, no sessions
Nobody starts a new session every time you text them, so she doesn't either. One window holds the whole stream: 150k tokens by default, up to 500k. When it reaches four fifths, a handoff runs in the background. The older past becomes a short recap, the newest fifth stays word for word, and she keeps talking while the recap is written.
She plays Minecraft herself
I didn't want a bot that plays while she comments on it. The game tools are hers, called by the same mind that answers on Discord. 27 of the 37 take time and run beside her, so she keeps talking while she mines or crafts. Reflexes like eating, fighting back, fleeing and landing a fall live in the mod and answer in game ticks: a fight wakes her, a snack doesn't.
The mod is BeaCraft, a client-side Fabric mod I wrote. The server sees an ordinary player, so it works on any vanilla server.
One session on 30 September: 46 minutes on a vanilla server, every game action her own decision. One block is one tool call.
- craft_item37
- find_block14
- attack_entity12
- smelt_item6
- move_to5
- mine_block4
- go_to_surface3
- build2
- move_away2
- 5 more tools5
- game actions
- 90
- lines spoken
- 158
- of 7.06M prompt tokens served from the cache
- 91%
A memory of her own
Everything she remembers lives in one SQLite file, with embeddings computed on the machine. After a session she writes a diary entry about it, and on every turn the entries that resemble what is happening come back into her prompt, split by who said them. She makes things up on purpose, so her own past lines come back marked as hers, never as facts.
Everyone she meets gets a tally. Only the people who matter, someone who donated, talked with her one to one or whom she chose to remember, get a card with their names across platforms, what she knows about them and how she feels about them. At night she dreams: the day's sessions become lasting memory and what she has learned about herself.
A voice in the call
In Discord calls a Node.js bot captures the room and plays her voice back. I wanted a call with her to feel like a call, not a walkie-talkie. You can talk over her and she lowers her voice, then stops. If you pause in the middle of a sentence, she waits for the rest before answering. And she starts saying a line the moment its closing quote arrives, without waiting for the rest of her answer.
- to her first sound in the call
- 1,226 ms
- model round trip, after the first call
- ~300 ms
- to lower her voice when you talk over her
- 280 ms
A face on stream
On stream she needs a face, and streamers use different kinds, so there are three: a PNG per mood, a 3D VRM model, or Live2D through VTube Studio. She breathes, blinks, nods while you talk and looks away while she thinks, and in calls her mouth follows what the room is actually hearing.
Yours to run
I wanted anyone to be able to run her. It is open source and self-hosted. Eight model providers sit behind one interface, in pools per role that rotate and fall back when one fails, and one of them is your own computer: with Ollama or LM Studio, a local whisper and Kokoro, she runs without a single API key. A web skill lets her search and read pages; it is off by default and refuses private addresses. Updating merges your edits to her prompts instead of overwriting them, and a control room in React and FastAPI shows everything she is doing.
Gallery





Want something like this?
I can build the same thing for your product.
Tell me what you have in mind. You get an answer the same day, with a realistic scope and timeline, not a sales pitch.