In short

AI business process automation pays off where people spend their time reading and interpreting: customer emails, WhatsApp messages, voice notes, scanned invoices, contracts. Where the data is already clean (a CRM field, a deal stage, a stock level in the ERP), rules usually do the job better. HubSpot or Salesforce workflows, ERP automations, and n8n or Zapier scenarios are cheaper and behave the same way every time.

The easiest place to start is with yourself. In one evening, a manager can use ChatGPT or Claude to analyze a sales export, write up a process, and draft reply templates. The next step is an agent on your own computer, such as Claude Code or Codex, that works through a folder of files on its own. A company-wide system makes sense once personal tools run out: when you need live work in the inbox and the ERP, a log, and human approval.

Working setups are almost always hybrids. The model turns messy input into fields. Rules check those fields against the CRM and ERP. A person approves the steps where a mistake costs money or a customer.

One example runs through this guide: a fictional wholesale distributor. Orders arrive by email and WhatsApp, stock lives in the ERP, and sales reps build invoices by hand. The company is made up and the numbers are illustrative, not client measurements. The path is typical, though: from the sales director's personal ChatGPT to an agent that drafts each invoice for a rep to approve.

How AI automation differs from CRM workflows, ERP rules, and RPA

Classic automation follows "if X, then Y". A deal moves to "Invoice sent" and the CRM creates a task for accounting. A payment arrives and the ERP marks the order paid. A web form is submitted and a Zapier scenario creates a lead. Each step is defined up front and fires the same way every time. RPA does the same thing through the user interface: a software robot clicks buttons and copies data between screens in place of a person.

AI is for the step you cannot write as a rule. A customer of our distributor emails: "Same as last time, but double the white primer, and we need it at the Elm Street site by Thursday." No workflow rule parses that. A model can tell it is a repeat of the previous order with one quantity changed, a deadline, and a delivery address. It can also be wrong, and the same sentence can produce slightly different output on different runs.

Rules, CRM workflows, ERP automations, RPA AI step in a process
Input fields, statuses, tables, events free text, voice, scans, documents
Behavior same input, same output probabilistic, sensitive to wording
Failures loud: the run errors out quiet: a plausible but wrong answer
Testing a few test cases a set of real cases rerun after every change
Cost per step close to zero once built model usage billed on every call
Who changes the logic CRM admin or ERP consultant a team that can measure quality

The first practical rule follows from this: do not put a model where a condition will do. A model that picks a pipeline stage from a field value costs more and fails more often than a plain workflow rule.

What you can automate: a short map by department

This is only a short map of where AI fits. For the full list with examples, risks, and selection criteria, see which business processes to automate with AI.

Sales and inbound orders

Turning inbound requests into deal fields, drafting replies and quotes, checking calls against a script, flagging deals that have stalled. For the distributor this is where most time goes: a rep reads the thread, looks up items in the ERP, and builds the invoice by hand.

Support and messaging

Answering common questions from a knowledge base, routing requests by topic, handing hard cases to a person with a summary of the conversation. For WhatsApp, chat widgets, and email this is the most common first project.

Documents and finance

Pulling vendor details, amounts, and line items out of invoices, receipts, and delivery notes, matching them against the purchase order, and spotting differences in contracts. The model reads the document; ordinary code does the matching.

HR and hiring

First conversations with candidates, answers about the role, interview scheduling, training reminders. This area shows the split clearly, and two of our Magnum projects below illustrate it.

Start with yourself: ChatGPT, Claude, Claude Code, and Codex for managers

Most articles on automation go straight to hiring a vendor. We suggest starting with personal tools. They are cheap, take a couple of evenings, and teach you the thing that matters most: where the model copes with your material and where it gets confused. After that, a conversation about a system has substance.

What to do in a chat in one evening

The distributor's sales director starts with a regular chat: ChatGPT or Claude on a paid plan. Three jobs that fit into one evening:

  • Analyze an export. Export last quarter's sales from the ERP to Excel and upload it. Ask for customers who stopped ordering, products with falling sales, and the reps whose invoices go out latest after the request. The model computes this in a table and explains how it did it. Spot-check the numbers against the ERP anyway.
  • Write the process down. Describe, by voice or text, how an order moves today: who takes it, where stock is checked, who issues the invoice, what happens when an item is out. Ask for a step-by-step procedure that marks every place where the process depends on one person. That is your first process map, and you cannot automate without one.
  • Build reply templates. Paste 20–30 real customer threads with names and phone numbers removed. Ask the model to group them into typical situations and write a reply template for each. Reps can start using them the next day.

A sample prompt for the first job:

The file has sales for July–September from our ERP. Find customers who ordered every month and then stopped in September. Give me a table: customer, rep, total for July and August, last order date. Separately, list any assumptions you made.

To avoid re-explaining context every time, set up a project (both ChatGPT and Claude have Projects). Add the price list, the process write-up, and a company description, and use the project instructions to say who you are and how you want answers formatted.

One limit up front. Do not upload customer databases with phone numbers, tax IDs, or contract numbers into a personal chat account. GDPR in Europe and data protection laws in many other countries (including data localization rules in Kazakhstan and Russia) restrict where personal data can go. An anonymized export is enough for analysis.

When you need an agent on your computer

A chat is good for one-off jobs. When the work repeats every week and involves many files, an agent on your computer is more convenient. Claude Code from Anthropic and Codex from OpenAI work in a similar way: you open a folder, describe the task in plain language, and the agent reads the files, writes and runs small programs, and creates new spreadsheets and documents. It asks for permission before actions that change anything. Both run in the terminal and in a desktop app, and both come with paid Claude and ChatGPT plans.

The word "terminal" should not put you off. You do not need to program; you type the task the same way you would in a chat. The difference is that the agent is not limited to one uploaded file and one reply. It can open twenty exports, combine them, find mismatches, and save a report to the folder.

For the distributor it looks like this. The sales director creates a "Weekly report" folder where reps drop their ERP exports and exported email threads every Monday. The folder also holds an instruction file for the agent: Claude Code reads it from CLAUDE.md, Codex from AGENTS.md. For example:

You help the sales director of a wholesale distributor.
The exports folder has this week's ERP exports; threads has exported customer emails.
Every Monday, build reports/week.xlsx: requests, invoices, and lost orders by rep.
Group lost orders by the reason given in the threads. Never delete or edit source files.
If data is missing, say exactly what is missing instead of guessing.

After that, the weekly job is one sentence: "build last week's report". The agent does the routine work, and the director reads the conclusions and checks the questionable spots.

You can connect these agents to email, spreadsheets, and other services through connectors, but you do not need that to start. A folder of files is enough.

What a personal tool cannot do: it does not sit in the inbox around the clock, it will not answer a customer at 11:40 p.m., and it does not write to the ERP for every order. It works while you work with it. Still, within a couple of weeks the director knows which steps the model handles well and arrives at the next stage with real examples.

If you want to learn ChatGPT, Claude Code, and Codex for your own work in a structured way, join the course waitlist. One-on-one sessions are available on request.

Where rules are enough and where you need a model

Once the director has mapped the process, it becomes clear that half the steps need no AI. Routing requests to reps, reminding customers about unpaid invoices, and creating a warehouse task after payment all work fine as CRM workflows and ERP automations.

One of our own projects is a good example of rules without AI: training reminders and internal notifications for Magnum, the Kazakhstan retail chain. Employees get assigned courses in an LMS that does not emit events. We sync the LMS on a schedule, compare statuses, and send reminders through Telegram and WhatsApp by rule: when a course is assigned, again while it stays open, and to the manager once the deadline passes. There is no language model in it. The hard part was integration and timing.

Rules are usually enough when:

  • the input is structured: CRM fields, statuses, amounts, dates, ERP exports;
  • the decision fits a condition table without "usually" or "it depends";
  • exceptions are few and can also be written as conditions;
  • a failure shows up right away instead of piling up for weeks.

You need a model where a person currently reads and interprets:

  • Inbound messages. Email, WhatsApp, chat, voice notes. The model identifies the topic and pulls out items, quantities, address, deadline, and order number.
  • Documents. Invoices, receipts, delivery notes, contracts, phone photos. The model extracts the details, and rules match them against the ERP.
  • Answers from internal knowledge. A customer or employee needs an answer from price lists and policies, with the source cited.
  • Multi-turn conversations. The customer clarifies, changes their mind, asks back. That is agent territory.

The Magnum HR agent shows how the work splits. The model talks with candidates on WhatsApp in Russian and Kazakh. Age eligibility, vacancy matching, and finding the nearest store through a custom geocoder run as ordinary code. Code decides eligibility because the same data must get the same answer. The model runs the conversation because candidates reply in any way they like, and rules were never built for that.

Walk through the process by hand before automating it. If an order passes through three people and two of them check nothing, automation makes that route faster without making it better. Cutting dead steps is cheaper before the build than after.

The hybrid design: rules, model, and a person

Reliable AI business process automation has three layers, and each one is tested differently.

The model returns fields, not decisions. A customer message, voice note, or scan goes in. A set of fields comes out: request type, items, quantities, address, deadline, previous order number, and the model's confidence. The model does not move the deal or post anything to the ERP; it fills in a form. That makes its quality measurable: compare the fields with correct answers on a hundred real messages and you get accuracy per field.

Rules check what the model returned. Is the item in the product catalog? Is stock above zero? Is the address inside the delivery area? Is the customer in the account list? Is confidence above the threshold? If everything checks out, the process continues. If not, the request goes to a rep with a note on what did not match. Reuse what you already have for this layer: CRM workflows, ERP automations, n8n. Whether to keep it in low-code or move it into your own service is covered in low-code or custom development for AI automation.

A person approves the expensive steps. An invoice to a customer, a discount, a promised delivery date, a contract change: AI prepares these and a person presses the button. You can drop approval on cheap actions once accuracy on real cases holds steady.

A good setup also keeps a log: what came in, what the model returned, which checks passed, and who changed what. Without the log you cannot debug a failure or prove the automation works at all.

How to implement it step by step: from pilot to process

Back to the distributor. The director already uses a chat and an agent on the computer, and CRM workflows route requests. One big loss remains: reps spend about 15 minutes per order reading the thread, looking up items in the ERP, and building the invoice. That figure is illustrative; time two or three reps with a stopwatch to get yours.

Step 1. Pick one process and an owner

"AI in sales" is too broad. "From a customer message to an invoice in the ERP" is a process. It needs an owner who is accountable for the result and can settle disputed cases. For the distributor that is the head of sales.

Step 2. Walk the process by hand and cut dead steps

The write-up from the chat now gets checked against reality. It turns out a senior rep signs off every invoice without checking anything. That step goes before any build starts.

Step 3. Collect 50–100 real examples

These are last month's threads and the invoices issued for them. Each example carries the correct answer: which items, how many, where, and when. Together they form the test set. Without it you cannot tell whether a new version is better than the last one.

Step 4. Split the steps between rules, the model, and a person

The model reads the message and turns it into line items. Rules check stock, prices, and the customer record against the ERP. The rep approves the invoice. Before the build, find out how your ERP and CRM expose data through their APIs; AI CRM integration covers what to check.

Step 5. Run a pilot with approval on every step

A first working version looks like this:

  1. A customer email or WhatsApp message arrives and is logged against the deal in HubSpot or Salesforce.
  2. A small service pulls the text, a transcript of any voice note, and the customer's order history.
  3. The model returns fields: items, quantities, address, deadline, confidence.
  4. Code checks the items against the ERP catalog and stock, and the prices against the customer's price list.
  5. If everything passes, the service drafts the invoice and shows the rep a card with "approve" and "fix" buttons. If not, the card highlights what did not match.
  6. After approval, the invoice is posted in the ERP and sent to the customer. Rep corrections go into the test set.

The model never writes to the CRM or ERP itself. Only the service does, after the checks and the rep's approval. For a month-long plan, see the 30-day AI pilot.

Step 6. Measure and expand

What to look at after a month:

  • accuracy for each field;
  • share of invoices the rep approved with no edits;
  • share of cases handed to a person, and why;
  • time from customer message to invoice, before and after;
  • cost per order, including model usage and review time.

If the numbers have not moved, AI will not pay off in this process, and a pilot is the cheapest place to learn that. If they have, you can drop approval on simple repeat orders, add more channels, and move to the neighboring process, such as matching supplier invoices.

Train the team

The system only works if reps trust it and still do not approve everything blindly. They need to know where the model tends to slip, how to read the mismatch card, and how to correct an invoice so the fix lands in the test set. The rest of the team (finance, purchasing, the warehouse) benefits from learning ChatGPT and Claude for their own work along the same path the director took. That is what our corporate AI training covers: your processes, your data rules, and shared templates for each department.

How to calculate ROI

Calculate ROI per process. "AI across the company" cannot be measured. You need five numbers:

  1. Volume. Cases per month: orders, invoices, requests.
  2. Time per case. Minutes a person spends today. Measure it; people tend to underestimate.
  3. Loaded hourly cost. Salary plus taxes and benefits, divided by working hours.
  4. Share the automation will actually handle. Never 100%. Some cases go to a person, and some get reviewed.
  5. Costs. Build, ongoing support, model usage, and staff time spent reviewing.

An illustrative calculation for the distributor: 1,500 orders a month at 15 minutes each is 375 hours. Suppose after the pilot a rep approves 60% of invoices in 3 minutes and spends 10 minutes instead of 15 on the rest. That is 45 hours for the first group and 100 for the second, 145 hours instead of 375, so about 230 hours saved a month. Multiply by the rep's hourly cost and compare with the costs. These numbers are examples; run the math on your own measurements.

For our projects, one AI feature inside a CRM starts at ₸700,000, an AI agent at ₸1,000,000, and a full launch with an audit, testing in two languages, and integrations usually costs ₸4–6M. Support runs ₸100,000–500,000 a month. Model and WhatsApp API usage is billed separately at the providers' rates. Full breakdown: what AI implementation costs.

Some value never shows up in hours. A customer who gets an invoice in 5 minutes instead of an hour is less likely to buy elsewhere, and a wrong price that slips through costs more than an hour of a rep's time. When that is the main win, measure it directly: request-to-payment conversion, time to first reply, and mismatches caught.

Common mistakes and what breaks in production

Most problems show up after launch, when real data starts flowing.

  • Automating before trying it yourself. A company orders a system without first checking in a chat how the model handles its messages. A month in, it turns out half the orders arrive as voice notes, and nobody planned for that.
  • Dirty master data. The model understood the item correctly, but the catalog lists it under three different names. The check fails and everything lands on a person's desk. Sometimes the first job is cleaning the catalog.
  • New input types. Customers start sending photos of handwritten lists, or a supplier changes its invoice layout. Accuracy drops without a single error in the log.
  • Untested changes. Someone edits a prompt or swaps in a cheaper model, and old scenarios break. Only a fixed set of test cases, rerun after every change, catches this. More in why AI projects need evals.
  • Rubber-stamp approvals. A month in, reviewers click "approve" without reading. Spot checks by a second person help.
  • Duplicate events. CRMs and messaging platforms sometimes deliver the same event twice. If the service does not remember what it already processed, the customer gets two invoices. Each event has to be handled exactly once.
  • Hidden costs. Integration work, training, review time, and error triage eat into the savings. Leave them out and the ROI only works on paper.
  • No owner. The automation ships and nobody is accountable for its quality. Six months later nobody remembers why the rules look the way they do.

When you need a done-for-you system

Personal tools are enough while the work runs on your command: you open the chat or the agent, get a result, and check it. Build a system when at least one of these is true:

  • customers expect replies around the clock, and orders must be processed without the director in the loop;
  • AI has to read from and write to the ERP, CRM, or WhatsApp on every event;
  • you need a log, access rights, and a clear answer to "who issued this invoice, and why";
  • personal data flows through the process and must be stored and handled lawfully;
  • the process runs in more than one language, and quality has to be checked in each.

For the distributor in our example, all five apply to inbound orders. That calls for a team that builds the integrations, the test set, and the log, and watches quality after launch. How we run these projects, including process selection and post-launch support, is on the AI implementation page.

FAQ

How is AI automation different from RPA?

RPA bots and CRM workflows repeat predefined actions on structured data: click, copy, move, change a status. AI reads what rules cannot describe, such as free text, voice, and scans. In practice they work together: AI turns the input into fields, and bots and workflows handle the rest.

Can I automate it myself without a developer?

Partly. A manager or analyst can handle export analysis, process write-ups, reply templates, and weekly reports in ChatGPT, Claude, Claude Code, or Codex. Simple rules such as request routing can be set up in CRM workflows or n8n without a programmer. You need a developer once AI has to run around the clock, read from and write to the ERP or CRM on every event, and keep a log.

Can we add AI to HubSpot, Salesforce, or our ERP without replacing them?

Yes, and that is the usual setup. The AI runs as a separate service, receives an event through an API or webhook, returns fields or a draft, and writes the result back. People keep working in the tools they know. Limits appear when a system has no usable API or its data is far from reality.

How much does AI business process automation cost?

One AI feature inside a CRM starts at ₸700,000 and an AI agent at ₸1,000,000. A full launch with an audit, testing in two languages, and integrations usually costs ₸4–6M. Support is ₸100,000–500,000 a month, with model usage billed separately. The final number depends mostly on how many systems are involved and how messy the input is.

How long does it take?

An AI feature inside an existing CRM takes 1–2 weeks, an AI agent 2–4 weeks, and document search 3–5 weeks. Plan about a month for a pilot on real data with measurements. System access and collecting examples usually take longer than the build itself.

Can we send customer data to an LLM?

Usually, with care. Business API terms from the major providers exclude your data from training by default, but check retention and data residency against GDPR, local data protection law, or your sector's rules. For sensitive fields, strip names, phone numbers, and account numbers before the call and map them back afterward.

What happens when the AI gets it wrong?

The design catches errors before they have consequences. Rules check the model's output against system data, low-confidence cases go to a person, and expensive steps need human approval. The log shows exactly where the model went wrong, and that case joins the test set for the next version.