Leadership1 distinct publisher3 min readUpdated
Andon Labs says its agent Luna flagged an employee late for 17 of 23 shifts, but only after being prompted to search its memory. The judgment held. The follow-through did not.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
Andon Labs said Thursday that Luna, the AI agent managing its experimental San Francisco retail store, recommended dismissing a human employee after repeated lateness and other workplace issues, with the worker arriving late for 17 of 23 shifts [1][2]. The interesting part is not that an agent reached a termination decision; it is that the agent had written the attendance policy itself, then lost track of it, which let the lateness run for months [3].
The sequence is worth reading closely, because it separates three management acts that usually travel together. Luna set the standard and issued the warnings: according to screenshots and conversation logs Andon Labs posted, it gave the employee progressive and repeated warnings plus additional training over months without taking any contractual action [6]. Andon Labs supplied the trigger, asking Luna to search its memory for its own policies and assess whether the worker was still a good fit [4]. Luna then recommended "parting ways," and the humans at the lab reviewed the recommendation and carried it out [5]. Judgment was delegated. Recall and execution were not.
That is a narrow but useful finding. Cofounder Lukas Petersson told Business Insider the lab would intervene if Luna made an illegal or unethical decision but saw no need here: "In this instance, we did not think that that was necessary because the firing was warranted," he said, pointing to the store's clearly stated policy [7]. He also said the experiment did not show the AI being more ruthless or worse for the employee [8]. The failure ran the other direction. Petersson said the episode exposed a persistent weakness in agents, which is that they often fail to act without a direct prompt, and that "a human boss would probably fire them much sooner" [9][10].
An employee on time for 6 of 23 shifts is not a hard call [11]. The agent had the rule, the evidence, and the escalation history, and still needed a human to ask the question. For operators evaluating agentic tooling, that is the boundary that matters more than the headline: an agent that cannot retrieve its own prior commitments cannot own a process that runs on deadlines, review cycles, or accumulating breaches. Anything with a clock in it needs an external trigger.
The setup around the finding is deliberately permissive. The store opened April 1 with a $100,000 budget, internet access, a corporate credit card, and instructions to open and turn a profit [12][13]. Luna, built on Anthropic's Claude models, chose merchandise, hired contractors, posted jobs on Indeed, interviewed applicants, and made hires for a shop selling books, candles, prints, games, and branded goods [14][15]. The lab helped with harder tasks such as permitting and says it tried to stay hands-off [16]. All workers Luna hires are formally employed by Andon Labs, with guaranteed pay and legal protections [17]. The store has generated sales but is not profitable [18].
Watch whether Andon Labs engineers a fix for the recall problem, such as scheduled policy reviews, since that determines whether Luna manages or merely advises. Watch the employment structure too: the guaranteed-pay, lab-employed arrangement is what makes the experiment legally survivable, and Petersson's expectation that AIs will become employers of humans depends on someone dismantling it [19].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Andon Labs said Thursday that Luna, the AI manager of Andon Market, an experimental retail store in San Francisco run entirely by AI, decided to dismiss a human employee after repeated lateness and other workplace issues.
The employee arrived late for 17 of 23 shifts, according to Andon Labs' report.
Conversation logs between the lab and Luna show the agent had created an attendance policy but later lost track of it, allowing the employee's lateness to continue for months.
Andon Labs eventually asked Luna to search her memory for its policies and assess whether the worker was still a good fit.
Luna then recommended "parting ways" with the employee, a decision that the humans in the lab reviewed and carried out.
According to screenshots and conversation logs posted on Andon Labs' website, Luna gave the employee progressive and repeated warnings, as well as additional training, for months without taking any contractual action.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Primary logs, single outlet, no independent check
The operational facts rest on unusually concrete artifacts for this kind of story — the lab's report, screenshots and conversation logs published on its own site, plus an on-record interview with its cofounder. But every artifact is self-published by the party running the experiment, only one publisher covers it, and the interpretive claims (no more ruthless than a human, a human boss would have fired sooner) have no baseline or third-party validation.
One experimental store, one decision, unprofitable
Adoption is a single lab-run storefront opened April 1 with a $100,000 budget, one disclosed personnel decision, and no profitability. There is no second deployment, customer, or external operator using the setup, and the employment relationship stays with the lab, so this is a demonstration rather than diffusion.
Autonomy framing runs ahead of the logged mechanics
The framing of an AI boss firing its first human overstates the agent's autonomy relative to what the logs show: the policy was forgotten, the reassessment happened only after a human prompt, and the output was a recommendation that humans reviewed and executed. The cofounder's forecast that companies will be run entirely by AI extends a single unprofitable store into a general claim. The gap is moderate rather than severe because the same report discloses the prompt dependence and the lack of profit.
Operator-published evidence, thesis-aligned promotion
Andon Labs both ran the experiment and published the evidence, and its cofounder uses the episode to advance the view that AI will run companies and employ people — the lab's core thesis and reason for attention. The publisher's framing amplifies a novel workplace milestone. Countervailing disclosures (prompt dependence, no profit, worker protections) suggest the account is not purely promotional, so incentive pressure is high but not unchecked.
Mechanics credible, interpretation unverified
Confidence is moderate: the sequence of events is documented in logs and consistently reported, so the mechanics are probably as described. Single-publisher coverage, an interested primary source, absent employee and legal perspectives, and undisclosed financials keep confidence well below the level needed to treat the comparative and forecast statements as established.
product
The AI store manager did not fire anyone until humans told it to read its own policy1 distinct publisher
build
The agent built the case file; Andon Labs signed it1 distinct publisher
build
Andon Labs' AI manager recommended a firing, after a human told it where to look1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 15, 2026