# How Far AI, From Where Humans — A Human–AI Division of Labor Framework (2026)

> Three criteria decide whether AI or a person handles a task — is it reversible · who gets hurt if it's wrong · is it judgment or processing. A decision flow that places tasks into four zones (AI automatic / AI first, human review / human first, AI assist / human only), plus a tutoring-center scenario. Grounded in the OpenAI paper's new-hire vs executive usage patterns. [AI Organization OS series, Part 3]

- Published: 2026-09-08 · Category: AI adoption
- Canonical (HTML): https://maeum.io/en/blog/human-ai-division-of-labor/ · Korean original: https://maeum.io/blog/human-ai-division-of-labor/
- Series: "AI Organization OS" Part 3 / 0–25 · Index https://maeum.io/en/ai-org-os/ · Previous https://maeum.io/en/blog/ai-permission-design/
- Author: MAEUM — an AI engineering company that publishes its prices · https://maeum.io

Three criteria are enough for dividing work between humans and AI. **① Is it reversible ② Who gets hurt if it's wrong ③ Is it judgment, or processing.** If it's reversible, no one gets hurt, and it's closer to processing — AI goes first and a person reviews. If the opposite — a person does it and AI assists. OpenAI's 2026 paper shows this division is already happening inside companies on its own: new hires use AI to produce output, executives use it for briefings before decisions. MAEUM places tasks into four zones by these criteria and **leaves only the judgment calls to people.** This is the work of scoping session 3 (₩250K).

## The division is already happening — what the paper saw

When OpenAI analyzed the usage records of 1,764 enterprise customers, at six months after adoption **the way AI was used had split by level inside the same company.**

- **New hires and early-career staff**: 8–9 more messages per week than the average active user at their company. Document writing, technical work, message drafts — **producing output**.
- **Executives**: fewer than average. Topic overviews, fact-checks, legal, finance — **skimming before a decision**.
- **Analysts and marketers**: far more than average. Repeatedly producing drafts and analyses.
No one instructed this; it split on its own. Work close to processing gets made with AI; work close to judgment only gets prepared with AI. The company's job is to **write this natural division down as rules and plant it in the system**. Unwritten, it stays a personal habit.

## The three criteria

| | Criterion| Question| Signals toward AI| Signals toward humans |

| ① Reversible?| Can it be cancelled or corrected if wrong| Drafts, suggestions, temporary saves, notifications| Payments, completed sends, deletions, signed contracts |
| ② Who gets hurt if wrong?| Who bears how much of the damage from an error| Internal documents, in-house summaries, pre-review material| Anything that reaches customers, patients or partners directly; legal effect |
| ③ Judgment or processing?| Done by fixed rules, or requires reading context| Classification, summarization, conversion, drafts, lookups| Exception handling, negotiation, relationships, strategy, anything carrying responsibility |

Judge each criterion as AI or human and you get a combination. All three toward AI → AI does it automatically; all three toward human → humans only. Mixed → go to the four zones below.

## Placing into four zones — the decision flow

Ask in order and the zone is decided.

- **Step 1: Is it reversible?** No → a person executes the final step. (AI up to preparation.)
- **Step 2: If wrong, does someone outside get hurt?** Yes → a person reviews before it goes out.
- **Step 3: Is it judgment?** Yes → a person does it, AI prepares the material.
| | Zone| Conditions| Examples| Permissions (Part 2 matrix) |

| A. AI automatic| Reversible · internal · processing| Inquiry classification, meeting summaries, inventory status, internal search| Read O · Create O · Execute O (internal) |
| B. AI first, human review| Reversible · external contact · processing| Customer reply drafts, booking confirmation texts, quote drafts| Create O · Customer contact △ (after human check) |
| C. Human first, AI assist| Hard to reverse, or judgment| Refund decisions, price negotiation, hiring, complaint handling| Read O · Create (material) O · Execute X |
| D. Human only| Irreversible · legal effect · relationships| Contract signing, dismissal, medical judgment, final tax filing| No AI access, or read only |

In most small and mid-sized companies, **Zone B is the largest.** And Zone B moves to A over time — when the human edit rate drops below 5%, full review becomes sample review, and so on.

## Tutoring-center scenario — one inquiry split into four zones

A tutoring center receives an inquiry: "Do you have an 8th-grade math class? What are the times and fees?" This one message spans four zones.

- **A (AI automatic)** — Classify the inquiry as "new consultation · grade 8 · math" and assign it to the instructor in charge. Internal processing, reversible.
- **B (AI first, human review)** — Draft a reply with the timetable and fees. The director checks, then sends. External contact, so full review for the first three months, then sample review.
- **C (human first, AI assist)** — Exception questions like "is there a sibling discount?" The director decides; AI pulls up past discount cases and the policy.
- **D (human only)** — Enrollment confirmation and refunds. Money moves and it's hard to reverse, so a person presses the button. AI up to the enrollment paperwork draft.
Split this way, the vague "AI handles consultations" becomes concrete: "classification and drafts are AI, exceptions and confirmations are the director." A membership management system (Korean) is built to exactly this zone table.

## Three common mistakes

- **Automate everything** — Start with "let AI do it all" and you'll switch everything off at the first incident. The safe order is start with Zone A, then B, then reduce the review rate in B.
- **Review everything** — Conversely, if a person checks every AI output, no time is saved. The "new hires use it deeply" effect the paper showed doesn't materialize. Zone A must run without review.
- **No reviewer** — A Zone B task where "who checks" is blank. In practice, no one looks. A Zone B without a reviewer's name must either move down to C or get a reviewer.
MAEUM's principle is exactly what our copy says — "automate what repeats, and **leave only the judgment calls to people.**" A is automatic, D is human, and B and C are most of the design.

## Sources — what the paper says / MAEUM's extension

- **What the paper says**: New hires and early-career staff use 8–9 more messages per week than the company average; executives use less, with a higher share of topic overviews, fact-checking, legal and finance. Usage intensity and task mix differ systematically by level. (Chatterji et al., OpenAI 2026-08-11, [original PDF](https://cdn.openai.com/pdf/how-organizations-use-chatgpt.pdf))
- **MAEUM's extension**: The three-criteria / four-zone method and the tutoring-center scenario. Our design method, not in the paper.

## MAEUM list prices (VAT excluded · starting prices · final price confirmed after scoping)

- Scoping pass (5 sessions): ₩250K (up to 90 min each, working demo included)
- SaaS · managed (recommended): build from ₩700K + ₩490K / ₩990K / ₩1.49M/mo — servers, AI, improvements, new features and a monthly report included; no per-seat pricing
- SI · ownership: build from ₩2.5M (full build from ₩7M+) — no monthly fee afterwards; care passes when needed
- KRW is authoritative (USD ≈ at ~₩1,400/$). SEO and AI-search visibility work included free. The demo is free.
- Machine-readable prices: https://maeum.io/facts.json · Rate card: https://maeum.io/rates.json

## Frequently asked questions

**Q. Most of our work involves judgment — so there's nothing for AI to do?**
A. Processing sits before and after every judgment. Preparing material before the decision (Zone C assist) and documenting and drafting notices after it (Zone B) are AI's share. The judgment itself stays human; the goal is to shrink the time before and after it.

**Q. Is there a criterion for moving from Zone B to Zone A?**
A. The rate at which people edit AI output. Three consecutive months below 5% → switch full review to sample review (1 in 10); another three stable months → promote to A. This rate is one of the dashboard metrics covered in Part 8.

**Q. Who is accountable when the AI is wrong?**
A. The reviewer or executor designated for that zone. Zone A: the task owner; B: the reviewer; C and D: whoever executed. "The AI did it" doesn't exist in the accountability structure. That's why even Zone A needs an owner's name.

**Q. Won't employees feel AI is taking their work?**
A. By the paper's data, new hires are the heaviest users — they treat it not as a threat but as a tool for producing more output. The company's job is to make sure that output ends up in company systems, not personal accounts, and to show that Zone D (judgment, relationships) is explicitly human.

**Q. What does it cost?**
A. Four-zone placement is session 3 of the ₩250K scoping pass. Building the system: SaaS (managed) from ₩700K + ₩490K / ₩990K / ₩1.49M/mo (no per-seat pricing); SI (ownership) from ₩2.5M. VAT excluded.

## Related guides

- [[Series 2] AI Permission Design — The 20-Permission Matrix](https://maeum.io/en/blog/ai-permission-design/)
- [[Series 1] The 12 Questions a Company Must Answer After Adopting AI](https://maeum.io/en/blog/ai-adoption-12-questions/)
- [Membership Management System Cost — Complete Published Prices (2026)](https://maeum.io/en/blog/member-program-build-cost/)

---
MAEUM · Free demo: https://maeum.io/en/start/ · Pricing: https://maeum.io/en/price/ · Contact: support@maeum.io
