Personal agents and OpenClaw
Trying to get a team of agents to remember what I told them and finish the work. Giving them names was the easy part.
Malcolm Graham / learning AI in public
I’m Malcolm Graham. I’m building things with AI to find out what it can actually do, what it gets wrong, and how much of my life I’m willing to hand over to it. So far, that includes a team of agents and a video game starring my dachshund. I have questions about where all this is going. In the meantime, the dog has a game.
Currently
Published September 10 / updated September 13
Ruben now has thirteen chapters, a trip up the coast, and a concert finale. You can turn the sound off before starting, and puzzle progress survives a break. The next question is whether someone who didn’t build it can figure out what the hell to do.
Play it. Tell me where you get stuck.Trying to get a team of agents to remember what I told them and finish the work. Giving them names was the easy part.
When I find a site or tool that does something well, I want to know how it works and whether I can build something useful from the idea.
Notes
Who does what, what currently works, and what I’m still trying to get right.
I wanted my scattered information in one useful agenda. Getting it onto a phone raised some questions about what should ever leave my computer.
What an agent should be allowed to read, change, and publish, and how to recover when something goes wrong.
OpenClaw / agents
I want agents that can take a job and carry it through without making me explain the same thing every time. I also want to know what they can access and what they’re doing with it. A confident answer is easy to get. A finished job takes more checking.
Pancho coordinates the work. Doozer builds, Einstein researches, Patton handles operations, Leonardo looks at design, and Rosco checks security. Mo reviews accessibility. Thoreau is supposed to catch the moment the writing stops sounding like me. There are eight named roles across nine configured entries because the default entry also uses Pancho’s identity. This is the setup as checked on September 13, not a live view of what they’re doing.
Platform: OpenClaw and its Codex plugin moved from the June stack to 2026.9.4, through an intermediate 2026.6.35 smoke test. Node is now 26.8.2.
Execution: the Codex authentication problem was fixed, and a real agent request got a reply through ChatGPT/OAuth. The OpenAI API-key fallback was removed.
Routing: explicit Telegram ownership was added during the migration. Gateway and Telegram passed fresh probes on September 13; the upgrade’s deterministic agent reply test also passed.
State: the upgrade migrated configuration, shared SQLite state, agent databases, and existing transcripts. Historical task records remain distinct from current health.
These are the jobs I’ve assigned. Separate roles should make the work easier to check; otherwise I’ve just created more people to manage, minus the people.
Keeps track of the job, assigns the work, and brings decisions back to me.
Builds the requested change and tells me what changed. A small fix should stay a small fix.
Finds the evidence, checks competing explanations, and tells me what the sources actually support.
Looks after deployment, recovery instructions, and the maintenance that makes tomorrow’s work possible.
Checks whether the design makes sense on a phone, whether you can read it, and whether the buttons are usable.
Checks code, dependencies, credentials, and what might be exposed. I want the problem found before somebody else finds it.
Looks for the barriers that keep people from using what I’ve built. Access should be part of the design.
Checks whether the writing carries my meaning and sounds like something I’d say. Currently assigned GPT-5.6 Luna.
Nine configured entries, eight named roles, seven specialists. The default main entry and the separate pancho entry share Pancho’s identity; Telegram routes to main. Agent-to-agent access is enabled for the default entry and seven specialists. Configuration does not prove that every handoff has been tested.
This update comes from the actual configuration, installed plugins, health checks, and upgrade record. A feature appearing in a menu doesn’t tell me whether it works.
GPT-5.5 remains the default for main, Pancho, and six specialists. Thoreau uses GPT-5.6 Luna, which is also registered across the other entries. GPT-5.4 mini and several Claude models remain in the catalog; a catalog entry is not a tested fallback. No automatic default fallback is configured.
Codex, OpenAI, Telegram, browser, Canvas, and core memory plugins are enabled. The Codex reply path and Telegram transport have working test evidence. The local OpenClaw node service is now installed and running, with node pairing completed in the evening setup record. Host Desktop was still unavailable in that record, pending macOS Screen Sharing. Browser and Canvas enablement alone does not establish end-to-end readiness. Active Memory and the experimental CUA Computer plugin are disabled.
The 2026.9.4 release adds improved plugin and skill discovery, conversation-to-skill creation, GPT Image 2.5 support, cloud-session controls, and terminal questions. These are release capabilities to evaluate; this review did not test their use in this setup.
The next test is one small job carried through a specialist, Mo’s accessibility review, and Thoreau’s voice check. Record what each contributed and whether the result improved. Memory retrieval and a full backup restore also need testing; the upgrade still reported memory alignment and backup housekeeping issues.
Public evidence snapshot, reviewed . The successful reply test belongs to the completed upgrade; this website review rechecked configuration, Gateway, Telegram, and model authentication without launching new specialist jobs. Private prompts, credentials, messages, and account details are not part of this report.
Field guide / security
The more useful an agent gets, the more of my information it wants. Before I connect another account, I want to know what the job actually requires, what the agent can change, and how I can take that access away. Convenience is a poor reason to stop asking.
OWASP’s GenAI LLM Top 10 now has a 2026 edition; the original project page is an archive. NIST’s published AI RMF 1.0 remains available while a revision is underway, with a 2026 critical-infrastructure profile concept note. Use these current source pages alongside CISA’s Secure by Design guidance. Reviewed September 11, 2026.
Files, memory, and notes stay in the private workspace first.
Tools run with least authority and narrow task context.
Risky, public, costly, or account-level actions require review.
Tests, diffs, and audit checks happen before public deployment.
Only approved and verified changes leave the private workspace.
Store preferences and working notes where they belong. Before publishing, check for private messages, credentials, and other people’s information. A useful example does not need to contain somebody’s actual life.
Use existing sign-ins, a password manager, short-lived tokens, and permissions limited to the service and task. Keep raw passwords out of chat.
Permission to read something does not include permission to send it. Be explicit about spending, account changes, messages, and publishing. Reading and local edits still need care when private information is involved.
Read the changes, run the relevant tests, check for secrets, and open it on a phone. After publishing, check the actual site. A successful deployment message doesn’t tell you whether the page makes sense.
In place: local workspace memory, Pancho as coordinator, specialist handoffs, explicit cost limits, and conservative rules for external actions.
Still needs work: long-term memory curation should be kept current so old preferences do not drift away from how the system actually works.
Scale: I need controls I can understand and maintain. More paperwork would not make this system safer by itself.
Field notes / tools
Here are the projects and tools. Some answer a fairly ordinary question. One puts my dachshund in charge of a dragon problem. Each has a link so you can see what it actually does.
I made a game for my wife and put Ruben in charge. Thirteen chapters of sniffing, digging, dragons, and a considerably longer trip than a dachshund needs.
Play + project storyA visual map of Pancho, specialist agents, memory, permissions, and safe automation inside this OpenClaw workspace.
Live noteAdd up the hardware, subscriptions, connectors, and things you forgot you’d have to pay for. Change the numbers to match your setup.
Live toolWhat an agent should be allowed to read, change, and publish, and how to recover when something goes wrong.
Live guideA look at surveys and routed model traffic, with the dates and limits attached. Popularity is not a test result.
Live noteLooks for patterns in text and images inside your browser. It cannot prove that AI made something, however convincing the number looks.
Live toolA few questions about the work, the data, and who is responsible when it breaks. Your answers become a short plan you can edit.
Live toolRecurring jobs that might be worth automating, with the inputs and limits spelled out.
Live guideHow I tried to turn scattered information into an iPhone agenda, including the privacy problem caught before release.
Live storyLab notebook
Updated September 13, 2026 — game expansion, Agent OS progress, and LinkedIn writing link.
The record of what reached the site. Dates matter here: a working build in July does not tell you what works today.
Replaced the nearly black section backgrounds with blue and teal, added new geometric artwork, replaced the empty project grids, and gave the diagrams and tools more contrast. The site needed to look different without changing what it does.
Added blue and green accents, clearer geometry, and a look at Ruben’s world on the homepage. Fixed the game menu so it scrolls through all thirteen chapters, including the first three that were missing from the shortcuts. Existing saves still work.
Rewrote the personal copy, project notes, and tool explanations. Less distance, fewer slogans, more of the actual reasons I’m doing this. Also fixed the plan builder printing backslash characters instead of proper line breaks.
The game now reaches the coast and concert finale, with a silent-play option and persistent puzzle checkpoints. Homepage and article summaries now match the released build. I added a link to my LinkedIn writing on AI safety and surveillance.
Added the OpenClaw 2026.9.4 upgrade, the authentication and Telegram repairs, and the full agent roster. The page now says which features are configured and which have actually been tested. See the progress and next tests.
Refreshed the OpenRouter snapshot, clarified survey periods, updated the security references and agent roster, and separated archived project milestones from current release claims.
Ruben’s six-chapter browser adventure now has a public home, a project story, and a plan for playtesting, clearer controls, and deeper exploration. Read the article and play the game.
Moved the site to Human For Now. The tools and project stories came with it, and mjgivai.com still works.
Published the July Daily Agenda build story: getting Health and Messages connected, reaching TestFlight, and catching private data in a demo bundle before the next release.
Local site review now starts with a phone-accessible preview URL, not a desktop-only localhost link, because mobile is the first review surface for new MJGIVAI website material.
Added examples of recurring jobs worth considering, what each needs, and what a first version would produce.
Added the browser-only scorecard and short plan builder. The questions are about the work and who owns it, before anybody buys more software.
Production Turnstile uses the real domain widget, admin pages are kept out of crawler paths, and smoke-test comments were removed from the moderation queue.
Added the agent names and their jobs so the setup is easier to understand.
The calculator turns a personal-agent setup into editable one-time and monthly costs, then invites moderated discussion.
Daily Agenda / July 2026
My information was spread across email, meetings, notes, conversations, and health data. I wanted an agenda that could make sense of it. That also meant deciding how much personal information an app needed and what should stay on my computer.
The recorded July build is 0.2.0 (13): a SwiftUI dashboard, Apple Health sync, Messages and Outlook connections, and a pipeline running locally. That is the milestone documented here. A newer build or current TestFlight availability has not been verified.
Early TestFlight work exposed a demo shell where the production dashboard should have been. We replaced it with the real ContentView, then made archive and simulator checks part of the release habit.
HealthKit can return no data for a valid day. The fix was to treat no-data responses as zero, persist the last aggregate snapshot, and keep refresh behavior honest.
Full Disk Access and assistive permissions had to be handled before the pipeline could use local Messages context. The documented ingest failed softly instead of breaking the whole agenda.
The app briefly bundled a generated agenda snapshot with real personal context. The security review caught it, the resource was removed, and the next version shipped without private bundled data.
Usage signals / not a census
I wanted a better answer than whatever model is being shouted about this week. These are three different views: the latest Stack Overflow survey located in the September 11 review, an older AI-builder survey, and OpenRouter traffic through September 11, 2026. They measure different things. None is a census of everyone using AI.
These are separate survey measures, not slices of a whole. The first covers respondents; the second covers professional developers. Stack Overflow still lists 2025 as its latest published results when checked September 11, 2026.
Source: Stack Overflow 2025 AI surveyArtificial Analysis, H1 2025: 591 respondents to the model-family question. Multiple selections were allowed. Rechecked against the original report on September 11, 2026; a newer comparable edition was not located. These figures are historical consideration, not current production volume.
Source: OpenRouter, usage data through September 11, 2026; “This Week,” all models. Prompt and completion tokens, including reasoning; private traffic excluded. Model variants rank separately. Licensed under CC BY 4.0. This is a saved snapshot, not a live feed.
A survey response, a model someone considered, and a billion routed tokens are different measurements. Adding them together would give us a very confident piece of bullshit. Higher traffic does not establish better answers, more users, or lower cost. The source links are there so you can check the numbers and find newer data.
Client-side tool
Before adding AI to a job, explain the job. These questions cover the work, the data, and who will check the result. They produce a rough score and a short plan in your browser. There’s no account or API call, and a high score is not a guarantee that the idea is any good.
Choose one job and say who owns it, what improvement would count, who checks the result, and when to stop. If those answers are vague, the experiment is still vague.
Interactive note
Add up the setup bill and monthly costs. Every number is editable. These USD assumptions were reviewed September 11, 2026; they are not vendor quotes and exclude taxes and domain renewals after year one. MCP connectors are optional, and this calculator does not use them. Read the explanation on the standalone calculator page.
One-time setup
Monthly run-rate
Project / local analysis
This looks at repetition and sentence patterns in text, or metadata and pixel patterns in an image. It runs here in your browser; your sample is not uploaded. Those clues can be misleading. A person can sound like a machine, and a machine can sound like a person. This tool cannot settle an argument about who made something.
Use it to inspect patterns, not accuse somebody. The score is not proof.
Local results will appear here.
No server request is made for analysis.
The browser does the work with ordinary JavaScript.
Use the score as a lead, not a final judgment.
Writing / AI, surveillance & dignity
I write about what happens when technology’s idea of “safety” collides with the privacy of the person living with it. Here’s the article and the recent posts behind that conversation.
My LinkedIn profile ↗“Solve the specific problem a person has consented to solve using the least invasive technology reasonably capable of solving it.”
From license plate readers in Volusia County to cameras in residents’ bedrooms: where does useful technology cross the line?

These aren't gotchas. They're the questions every operator should be asking every vendor in this category, including mine.
Worth saying plainly: this was a lab study. Roughly 50 university students. A single session. Nobody has run this on an 84-year-old in a memory care room.
A resident moving into senior living shouldn't have to trade privacy and dignity for safety. And operators shouldn't have to choose between being innovative and being creepy.
About / contact
I like making things, figuring out how they work, and asking questions that don’t always make the sales presentation better. This site is where I’m working through AI by using it. Some projects are useful. Some are personal. I’m willing to be wrong, but I want to understand why, and I don’t want the machine polishing the opinion out of everything I say.