“Darling, I can't take this anymore. Mum's WhatsApp isn't working, Telegram won't connect, and VK calls turn her face into a blurry square creature. Do something, will you?”
“What am I supposed to do about it?”
“Think of something. You're a programmer now!”
“Errrrrr. Damn.”
Well, why not? I once had a six-digit ICQ number, you know. Practically a qualification. Let's make our own messenger! Send a message, receive a message, call, answer. What else do these things do? Emoji? GIFs? Please.
Actually, I lost that six-digit account because I forgot the password. Then I had a nine-digit one, 305128631. Still remember it! Or was I already using QIP by then? Never mind. For anyone spared this particular piece of internet archaeology: ICQ had numeric user IDs, and a short one carried a certain ridiculous prestige.
Everybody knows what a messenger is. How many are there? A thousand? Two thousand? Whatever the number, I should have researched which one worked in Russia and recommended it. So my wife and mother-in-law would have five more icons on their phones and poke around to find the one working that day. Inconvenient, but plenty of people did exactly that. Honestly, I should have talked my way out of it.
Unfortunately, I'm constitutionally incapable of leaving this sort of thing alone.
The decision shaped the next year of my life. I don't regret it at all. It was probably one of my most productive years: I don't think I'd ever learned so much, so quickly, across so many subjects.
I knew perfectly well I was reinventing the wheel. WeekendDiver had taught us to look at existing products before diving into code. We looked. Plenty of bicycles, each with something going for it, but none I particularly wanted to ride: awkward frame, handlebars pointing the wrong way, so many accessories you couldn't find the pedals.
What if we took a bicycle and strapped a rocket engine to it?
Back to the familiar approach: research first. What was there to research? T., W. and S., obviously. Those initials should thoroughly conceal the identities of the apps I'm about to complain about.
My questions were already rather different from those of someone who just wanted to call her mum. Looking at familiar apps, I'd catch myself thinking: that's good, that's convenient, and why is this such a mess? I'd do it differently.
A dangerous thought when there's an AI beside you that may answer “I'd do it differently” with “Let's do it now.”
With T., I liked the idea of a broad space for communication: private conversations, groups, channels, the whole circus. Cloud history and access from multiple devices are extremely convenient; I understand the attraction. I didn't like the service potentially having broad access to that history. Yes, there were separate secret chats, about which I had questions of my own. I wanted privacy to be the normal state of a personal conversation, without extra buttons to press. Fine, says the attentive reader, what about W.?
W. was a familiar, convenient everyday tool, but I didn't want to reproduce it under another icon. I was a developer now, damn it. I wanted private conversations, public channels the way I liked them in T., my own interface, a chance to change the irritating bits, another button here and one fewer over there. Somebody else's app has a predictable response to those wishes: none. Use what you're given.
I looked at S. with particular interest in its security design. Some decisions seemed excessive for the product I wanted to build. Seemed, given my understanding and needs at the time. I wasn't riding into the internet carrying a certificate saying “Audited All Cryptography Everywhere”. Just an old six-digit account I'd managed to lose.
We'll return to encryption separately; it gave me plenty to think about and rebuild. Here, the interesting part is how dissatisfaction became a design: keep what's useful, drop what's unnecessary, improve what's awkward, add what's missing. Preserve privacy where it matters, and don't burden the user with complexity just because I can now ask an agent to implement it.
Now that sounded worth trying. Why yes?
The family request to “make it possible to talk properly” grew into MangoConnect. I knew a tiny two-person talking app wouldn't hold my interest. If I was going to build a messenger, it would have proper chats, groups, calls, photos and all the things people expect.
“All the features.” A splendid way to start a project, particularly if you avoid counting how many features fit inside “all”.
First, though, we needed decisions that would shape everything else. A big one: what exactly were we promising people about their conversations?
I wanted personal conversations encrypted on participants' devices, with the server delivering them rather than reading them. I was already familiar with encryption, but hadn't appreciated how many everyday conveniences would need rethinking if we followed that principle consistently.
As a wish, it's simple: safe and convenient, preferably simultaneously, without extra questions, buttons or a textbook to read before your first message. Roughly the level of detail I began with.
I'm exaggerating. I do have a degree in protecting information against technical surveillance, with honours. Just so you don't picture a complete novice wandering cheerfully into the swamp.
Security came slightly ahead of convenience, though I wanted as much of both as possible. One decision was “one account, one active device”. I didn't want several additional connected clients continuously having access to the same conversation. That costs the user some convenience, so it had to be a conscious trade-off, not merely another item in a list of advantages.
I also didn't want a phone number to be compulsory for registration. Why did I need one if the job was helping people communicate? Even privacy-focused S. requires a number at registration, although you can hide it from the person you're talking to. I wanted to avoid that dependency and know rather less about the user in general.
Meanwhile, a public channel and a private conversation seemed like different things. Publishing to an open audience serves a different purpose, with different rules and requirements, from sending a message to your wife. I chose a hybrid design: personal conversations, closed groups, calls and group calls with end-to-end encryption; public groups and channels in a separate open mode.
I like versatile devices, as you may remember, but that doesn't mean treating every function identically. Different situations can coexist in one app if you understand their boundaries and can explain them to both users and agents.
The latter turned out to be a quest of its own. More on that later.
By this point, I could discuss application structure sensibly with agents: the pieces involved, what the server should do, what stayed on the phone, how they connected. The experience from ManiTalk and EasyStayTH hadn't evaporated. I'd actually been learning, so our conversations went beyond “draw me a screen and we'll see”.
Understanding an app in broad terms and anticipating every consequence of your wishes are rather different abilities, though. Especially when the wish is “an ordinary messenger, just with a bit of encryption, you know, DO NO MISTAKES”.
Take my cheerful “send a message, receive a message”. Two actions for the user, but plenty happens between them. The recipient may be offline, the app closed, the connection interrupted at the precise point where you can't tell whether the message went through. The user taps again because they want to send something, not investigate your connection state. All of this should feel completely ordinary: typed it, arrived. Nobody owes you admiration for the number of cases you've handled in between. Nobody cares, really.
Calls are similar. “Call, answer” sounds lovely until you describe what each phone must do before the answer, during the conversation and after someone unexpectedly loses their connection. We started because the family needed reliable calls, so “messages work, we'll get to calls eventually” would have been an odd result. I'd been told quite clearly that the work wouldn't be accepted without video calls.
We hadn't even reached the emoji I'd casually dismissed as easy. We were merely unpacking “messenger”. I knew these features as a user; I'd used them forever. Now those habits had to become development decisions. Reverse engineering, sort of, without the code.
The agent was useful here. I could describe the behaviour I imagined, it could break down what we'd need, and I could examine, clarify and change it. Gradually we had something concrete enough to discuss and then implement. Yes, I know GitHub and GitLab exist, with open-source repositories. But pride may be a deadly sin; the rewards are delicious.
The design brought more things to remember. We chose this for that reason, decided differently here, postponed that feature, and can't change this part independently of its neighbour. Yesterday, another agent and I understood all of it. Today, a new session made me the project's chief historian again.
A good time for the magic prompt from the previous article. Except we hadn't found it.
For a while, I was genuinely reluctant to start a new session because losing context hurt. It could happen during a session too, when “compacting…” appeared. At one point, my routine was to ask for a summary, open a new session, paste it in, check that the new agent seemed to understand, and only then close the old one. Ah, the early days.
With Mango, I was consciously trying to leave context inside the project. We began with a task list and an idea document, then added others. I didn't wake up one morning and draw a complete system. Something became necessary during work; we found somewhere to record it and a way to find it later.
Eventually we needed somewhere for agents to pass unfinished work to each other. “I tried something, it didn't work, but I nearly understand it” is valuable information, particularly if it survives the session closing. I'd rather not pay in time and usage limits for the next agent to reach the same “nearly”.
So the project that began with a request for family calls gradually acquired a memory of its own. My next job was getting the agents to use it instead of arriving as though they'd never been there before.
Building a messenger is easy, remember? All that remained was explaining what we were building to several very clever machines, and making sure they'd still know tomorrow.
Coming up: “Read the docs.” How I started moving the contents of my head into files, and why creating a documentation folder wasn't enough.
Sources and further reading:
