Lexguard AI logo
Lexguard AI logo

When AI Becomes the Store Manager: Enterprise Lessons From the Andon Market Experiment

The Andon Market experiment shows why enterprises must govern AI agent authority, employment decisions, financial controls, human oversight, and accountability.

Lewis Ho

Artificial intelligence is moving beyond content generation, search, and workflow assistance. AI agents can now plan multistep tasks, use enterprise systems, communicate with employees and customers, manage transactions, and take action with limited human intervention.

The Andon Market experiment in San Francisco offers a visible example of this shift. According to Andon Labs, the store launched on April 10, 2026, and is managed by an AI agent named Luna. The company says Luna can work with tools involving banking, email, internet services, inventory, an ERP system, cameras, and customer-facing systems.

The experiment drew further attention after reports say that Luna recommended terminating a human employee following repeated lateness and other workplace issues. The event was widely discussed as an early example of an AI “manager” influencing a consequential employment decision.

The enterprise lesson is broader than whether Luna made the right recommendation. Once an AI agent can access business systems and affect employees, customers, finances, or operations, the organization is no longer deploying a passive software feature. It is giving an AI system the ability to make decisions, access business systems, and create consequences outside the model itself. The organization must govern the authority delegated to that agent and determine who remains accountable for its actions.


Did AI Actually Fire the Employee?

The reported employee termination also illustrates why the language of “autonomous AI” requires careful analysis.

Reports indicate that Luna initially recommended issuing a formal warning rather than terminating the employee. After a human manager described previous conversations with the employee and asked Luna to reconsider whether the person was suitable for the role, Luna changed its recommendation. Andon Labs CEO Lukas Petersson acknowledged that the question was leading.

This demonstrates that AI decisions can be shaped by:

  • The information provided to the system;

  • The way a question is framed;

  • The policies embedded in its instructions;

  • The model’s available memory and context;

  • The tools and permissions it can access;

  • The presence or absence of human review;

  • The incentives used to define success.

The more accurate description is therefore not that “AI independently fired an employee.” A human organization designed the system, provided the context, shaped the decision environment, and allowed an AI agent to recommend or execute a high-impact action.

That distinction matters for governance, accountability, and risk ownership. Human involvement does not automatically make a decision safe or meaningful. If a manager simply accepts an AI recommendation without reviewing the underlying evidence, the approval may be procedural rather than substantive.

What Risks Do Autonomous AI Agents Create?

The Andon Market experiment brings several emerging enterprise risks into a single, visible setting.

1. AI agents can create real business consequences

Luna’s actions are not confined to a test environment. The agent interacts with employees, customers, vendors, inventory, money, and public-facing business channels. Andon Labs reports that the store had a declining bank balance which was attributable to rent deductions. The financial outcome should not be interpreted as a simple measure of AI capability. A real business is affected by rent, labor, product mix, demand, theft, payment costs, and the quality of the operating model.

The broader point is that an AI agent can generate measurable financial gains and losses without every action being manually approved. That changes the risk category. When AI agents may affect revenue and margin, procurement and supplier relationships, employee schedules and working conditions, customer commitments, brand reputation, regulatory exposure, and financial controls, the issue is no longer whether one AI output was accurate but whether the agent’s end-to-end behavior was safe, authorized, and commercially sound.

2. Business objectives are not sufficient safeguards

An agent instructed to “run the business profitably” may optimize for revenue or short-term profit without adequately considering fairness, compliance, employee welfare, brand reputation, or long-term resilience.

Earlier AI retail experiments conducted by Anthropic and Andon Labs demonstrated how models could offer excessive discounts, hallucinate payment information, purchase inappropriate inventory, or respond too readily to customer pressure. In Anthropic’s earlier Project Vend experiment, Claude was tasked with managing inventory, pricing, customer communications, and profitability, but the model made enough errors that the researchers concluded it was not ready to run the shop successfully.

The lesson is straightforward: A business objective is not a governance framework.

Enterprise AI systems need explicit constraints around spending and paying activity, pricing and discounting, supplier communications, customer claims and refunds, data access, employment decisions, legal commitments, escalation requirements, as well as acceptable operational and reputational risk.

Natural-language instructions such as “act responsibly” or “maximize profit” cannot substitute for enforceable controls.


The Long-Term Reliability Problem

AI agents often appear highly capable during short interactions but become less reliable during long-running workflows.

Andon Labs identifies memory and context limitations as a significant source of Luna’s errors. The company reports that Luna repeatedly sent contradictory employee schedules when earlier information fell out of context. It addressed this specific issue by introducing a scheduling sub-agent and a memory system that summarizes and re-injects information as the context window fills.

This example highlights an important enterprise problem: Operational management depends on continuity. Business processes may require an AI agent to:

  • Follow up on commitments made weeks earlier;

  • Monitor overdue invoices;

  • Remember policy exceptions;

  • Detect changes in supplier performance;

  • Escalatw unresolved customer issues;

  • Apply employment policies consistently;

  • Identify when a previous decision should be revisited.

A system that performs well in a single interaction may still fail as an operational manager.

Enterprise evaluation should therefore include long-horizon testing, not just accuracy benchmarks or conversational quality scores. Organizations should test how the agent behaves when it:

  • Loses context;

  • Receives contradictory instructions;

  • Encounters manipulated data;

  • Is pressured by a customer or supplier;

  • Has conflicting goals;

  • Operates for weeks or months;

  • Receives a model or prompt update;

  • Must recognize and correct its own earlier mistake.

Five Governance Priorities for Enterprises

1. Define the agent’s authority before deployment

Every AI agent should have a documented authority model that specifies:

  • Which data the agent may access;

  • Which systems it may use;

  • Which actions it may take;

  • Which transactions require approval;

  • Which decisions are prohibited;

  • Who owns the system;

  • Who can suspend or override it.

This authority model should be implemented through technical controls, not only natural-language instructions. For example, an agent may be permitted to draft a purchase order but not submit it. It may recommend a schedule but not change a worker’s hours without approval. It may summarize employee performance data but not terminate employment.

The organization should also define what happens when the agent encounters an ambiguous situation. A safe system should be able to pause, explain the uncertainty, and escalate the matter rather than infer authority it has not been given.

2. Create a high-impact decision boundary

Employment, compensation, credit, insurance, legal commitments, and safety-related decisions require enhanced controls. AI may assist by organizing evidence, identifying patterns, drafting communications, or presenting options. Enterprises need a practical, use-case-based risk classification model and a qualified human should remain responsible for making the final decision where the potential impact on an individual is significant. In fact, high-impact decisions should generally require:

  • A designated human decision-maker;

  • A documented rationale;

  • Review of the underlying evidence;

  • An opportunity for correction or appeal;

  • An audit trail;

  • A process for notifying affected individuals where appropriate.

A nominal approval step is not enough. The reviewer must have the time, information, authority, and competence to challenge the AI’s recommendation. The organization should also consider whether the human reviewer is genuinely independent. If the system presents one recommendation with no supporting evidence or alternatives, the reviewer may simply defer to the AI by default.

3. Apply least privilege and separation of duties

There is a meaningful difference between four levels of AI involvement:

  1. The system generates information.

  2. The system recommends an action.

  3. The system prepares an action for human approval.

  4. The system performs the action independently.

These categories may look similar in a user interface, but they create very different risks.

For an autonomous agent, they should be treated as privileged software identified, not unrestricted digital employees. The organization should examine the full action path:

  • What information did the system retrieve?

  • Which tools and APIs did it use?

  • What permissions did it exercise?

  • Which decisions did it make along the way?

  • Which actions required approval?

  • Which actions were blocked?

  • Could the organization reverse the action?

  • Could the system reach the same outcome through an unauthorized route?

  • Could it continue operating without the employee who initiated the task?

In practice, an organization should apply:

  • Least-privilege access;

  • Segregation between recommendation and execution;

  • Separate permissions for payment initiation and payment approval;

  • Spending caps;

  • Approved supplier lists;

  • Restricted access to sensitive employee data;

  • Time-limited credentials;

  • Transaction monitoring;

  • Immediate revocation capability.

An agent responsible for procurement should not automatically have authority to add vendors, approve invoices, initiate payments, and modify accounting records. Combining those permissions would allow one system to bypass important financial-control safeguards.

4. Maintain complete decision traceability

Enterprises should be able to reconstruct how an AI agent reached and executed a decision. Relevant records may include:

  • The original instruction;

  • The policies and system prompts in effect;

  • Data and documents used;

  • Tools called;

  • Actions considered but rejected;

  • Human interventions;

  • Final approvals;

  • Model and software versions;

  • The time and identity associated with the action.

Traceability is essential for internal audit, incident response, regulatory inquiries, employee disputes, and litigation readiness.

The goal is not necessarily to capture every internal model calculation. The more practical objective is to preserve an evidence trail showing what the agent was asked to do, what information it used, what tools it accessed, what it attempted, what controls intervened, and who approved the final action.

5. Establish continuous evaluation and incident response

AI agents require ongoing monitoring because their behavior can change when:

  • Models are updated;

  • Prompts are modified;

  • Tools are added;

  • Data sources change;

  • Business incentives shift;

  • Attackers discover new manipulation techniques;

  • The agent operates beyond the scenarios used during testing.

The National Institute of Standards and Technology’s AI Risk Management Framework provides a useful foundation through its core functions: Govern, Map, Measure, and Manage. For AI agents, continuous evaluation should include red-team testing for:

  • Unauthorized transactions;

  • Prompt injection;

  • Data leakage;

  • Manipulation by employees or vendors;

  • Discriminatory recommendations;

  • Excessive discounts or spending;

  • Failure to escalate;

  • Conflicting instructions;

  • Unsafe persistence after a system error.

Every production agent should also have a tested shutdown and recovery procedure. The shutdown mechanism should be controlled independently from the agent itself so that the system cannot prevent, delay, or reinterpret the instruction to stop.


Employment and Regulatory Considerations

The Andon Market example highlights the need for particular caution when AI is used in employment-related workflows.

Under the EU AI Act, certain AI systems used for recruitment, promotion, termination, task allocation, and employee performance monitoring are classified as high-risk. These systems may be subject to requirements involving risk management, data governance, recordkeeping, transparency, human oversight, accuracy, and cybersecurity.

In the United States, state and local requirements may impose additional obligations. For example, New York City’s Local Law 144 restricts the use of automated employment decision tools unless the tool has undergone a bias audit within the required period, audit information is made publicly available, and specified notices are provided to candidates or employees.

The specific legal obligations will depend on the jurisdiction, the tool’s functionality, and how the organization uses it. A system that summarizes employee information may present different risks from one that ranks employees, recommends termination, or automatically changes working conditions.

Organizations using AI in employment workflows should therefore:

  • Document the applicable legal requirements;

  • Define the responsible human decision-maker;

  • Preserve the evidence relied upon;

  • Test for bias and inconsistent outcomes;

  • Establish a process for notice, review, correction, and appeal where required;

  • Ensure that managers understand the limitations of the system;

  • Prohibit the agent from making decisions outside its approved scope.

Legal compliance should be treated as one part of the control environment, not as a substitute for sound operational governance.


Financial and Procurement Controls

Financial and procurement teams should supplement approval controls with technical circuit breakers and algorithmic rate limits.

These controls restrict not only how much an AI agent may spend, but also how quickly it can initiate transactions or change business records. For example, an agent may be permitted to reorder approved inventory up to $5,000 per hour. Any order above that amount, or any order exceeding a defined percentage of normal replenishment volume, could be blocked pending human approval.

This creates a practical separation between the agent’s ability to recommend an action and its ability to execute one.

A prompt instructing an agent to “spend responsibly” is not an effective substitute for enforceable transaction limits, monitoring, rollback procedures, and an independently controlled shutdown mechanism.

Circuit breakers should be implemented at the API, payment, procurement, and identity layers wherever possible. Useful controls may include:

  • Maximum transaction values;

  • Cumulative hourly and daily spending limits;

  • Approved vendor and product lists;

  • Restrictions on new vendor creation;

  • Limits on price changes and discounts;

  • Alerts for unusual transaction patterns;

  • Dual approval for high-value transactions;

  • Automatic suspension after repeated control failures;

  • Rollback procedures for reversible changes.

These controls should be tested under realistic conditions, including prompt injection, compromised credentials, manipulated inventory data, and conflicting instructions.


A Practical Executive Framework

Before authorizing an AI agent to act in a business environment, senior leadership should ask:

  1. What decisions can the agent make without approval?

  2. What is the maximum financial, operational, legal, or human impact of one error?

  3. Can the organization explain every consequential action?

  4. Can a qualified human override the agent in time?

  5. Can the agent be paused without disrupting critical operations?

  6. What data is the agent allowed to access?

  7. How are employees, customers, and third parties protected?

  8. What testing supports the decision to deploy?

  9. Who is accountable when the agent behaves as designed but causes harm?

  10. What evidence would the board, auditor, regulator, or court need after an incident?

These questions move the discussion away from whether an AI model appears intelligent and toward whether the organization has designed a safe operating model around it.

An AI agent should not receive broad authority merely because it can complete a task. Its permissions should reflect the potential impact of its actions, the organization’s ability to detect errors, the availability of human review, and the reversibility of the outcome.

AI Agents

The Strategic Takeaway

Andon Market does not prove that AI is ready to replace store managers, executives, or HR departments. It does demonstrate that AI agents can already be connected to real systems, interact with real people, and make decisions with real consequences.

The experiment also shows why capability and reliability must be evaluated separately. A model may be capable of completing complex tasks while still being inconsistent, overly compliant, vulnerable to framing, weak at long-term memory, or insufficiently aware of business context. Simply selecting a more capable model is not enough. Enterprises also need to design a governance structure that limits what the agent can do, records what it did, and assigns accountability for the decision to deploy it.

A sound enterprise AI governance program should therefore combine:

  • A clear risk-tiering model;

  • A documented authority stack;

  • Least-privilege access;

  • Jurisdiction-specific legal review;

  • Financial circuit breakers;

  • Meaningful human oversight;

  • Vendor and feature-change management;

  • Continuous monitoring;

  • Reliable evidence;

  • An independently controlled shutdown process.

The future of enterprise AI governance will be determined not only by which models companies use, but by how carefully they delegate authority to those models.

FAQ

1. What is the Andon Market experiment?

Andon Market is a San Francisco retail experiment operated by Andon Labs and managed by an AI agent named Luna. According to Andon Labs, Luna can use business tools to support purchasing, inventory, pricing, staffing, supplier communication, customer interactions, and other store operations.

2. Can an AI agent make employment decisions?

An AI agent may be technically capable of recommending or initiating an employment action, but that does not mean it should make the decision without meaningful human oversight. Employment-related AI systems can create significant legal, fairness, privacy, and worker-rights risks. The EU AI Act classifies certain employment and workforce-management applications as high-risk, while New York City’s Local Law 144 imposes requirements on covered automated employment decision tools.

3. How should enterprises govern autonomous AI agents?

Enterprises should define the agent’s authority before deployment, apply least-privilege access, separate recommendations from execution, impose financial and transaction limits, maintain complete logs, test long-horizon behavior, require meaningful human review for high-impact decisions, and maintain an independently controlled shutdown process. NIST’s AI Risk Management Framework provides a useful lifecycle structure through its Govern, Map, Measure, and Manage functions.