Cloudfinch
Back to Blog
AI6 min read

The Work That Wasn't Worth Automating a Year Ago

AI models got cheaper and more capable again this year. For an operating business, that moves the checking, reading, and chasing nobody has time for into the worth-automating column.

A cost-per-task curve falling below a dashed worth-automating line, beside a list of routine work now handled by agents: invoice line checks, intake documents, return reasons, and overnight reconciliation, with mismatches sent to a person for review

Cloudfinch Team

Oct 6, 2026

Most businesses have a list of work they know they should do and don't: checking every invoice line against its purchase order, reading every return reason, comparing every delivery with the window the customer was promised. Teams sample it, skim it, or leave it for month-end, because doing it for every record would take more hours than it saves.

A year ago, handing that list to AI often didn't pay either. The models could do parts of it, but the cost per record, the size of the files, and the length of the tasks got in the way.

Over the past twelve months the AI labs have moved all three, and a lot of that list is now worth automating.

What changed in twelve months

The releases came quickly this year. Five changes matter most for operating businesses.

  • The same performance costs far less. Epoch AI's September analysis finds that the cost of reaching a given level of AI performance has fallen about 47% per quarter since 2023, or roughly 13x a year.
  • The top models got cheaper. A year ago, Anthropic's top-tier model, Claude Opus 4.1, listed at $15 per million input tokens and $75 per million output tokens. Claude Opus 5.5 lists at $4 and $20 (Anthropic pricing, October 2026). Newer Claude models count about 30% more tokens for the same text, so the real saving is smaller than the list prices suggest. It is still large.
  • Mid-tier models now do top-tier work. Claude Sonnet 5.5, released September 28, scores two points below Opus 5.5 on GDPval-AA, a benchmark built from real professional tasks, at half the price. Anthropic still points to Opus for complex, open-ended work that needs sustained judgment.
  • Whole files fit in one request. Current Claude models read up to a million tokens at standard per-token pricing, and up to 600 PDF pages in a single request. That is enough for a client's full file or a contract with every amendment.
  • Agents can keep working. At DevDay on September 29, OpenAI launched Dots, always-on agents with their own cloud computer and connected apps. Anthropic bills the runtime of its managed agents at $0.08 per session-hour.
  • What a routine task costs now

    Anthropic's documentation includes a worked example: handling 10,000 support conversations, averaging about 3,700 tokens each, with Claude Haiku 4.5 costs about $37 in model fees (pricing docs). Work that can wait a few hours runs through the Batch API at half price, and repeated instructions served from the prompt cache cost a tenth of the normal input rate.

    For most routine work, the model bill is now a small line next to the hours the work used to take.

    The costs that remain are the ones that were always there: connecting your systems, cleaning the data the work depends on, and deciding who reviews what. They still deserve a careful plan, and they are where most of the effort in a good build goes.

    Work that just crossed the line

    None of these ideas are new. What has changed is that they pay at the volumes of a 10 to 500 person company.

  • Checking every record instead of a sample. Every invoice line against its purchase order and delivery note, every carrier bill against the rate card, every timesheet against the job it was billed to. The agent clears the matches and sends the mismatches to a person.
  • Reading every inbound document in full. Bills of lading, proofs of delivery, client intake packets, contracts with their amendments. With a large context window, the summary covers the whole file, including the detail on page 40.
  • Sorting the long tail. Return reasons, support messages, and every "other" typed into a free-text field, classified into categories someone can act on.
  • Overnight reconciliation. An agent matches the day's orders, shipments, and payments while the office is closed and leaves a short exception list for the morning.
  • A follow-up for every case. Payment reminders, delay notices, and missing-document requests drafted from the record for every customer, including the small accounts nobody had time for, and queued for approval.
  • Match the model to the step

    Cheaper models change how a system should be designed. A well-built workflow uses different models for different steps: a small, fast model to classify and extract, a mid-tier model to draft and reconcile, and the top model for the cases that need judgment. Anything the agent is unsure about goes to a person, with the context attached.

    This is how we design for a predictable monthly run cost. The routine bulk of the traffic goes to the cheaper models, caching and batching handle the repeats, and every proposal includes a projected monthly bill.

    Every release is an upgrade you can take

    The labs now ship meaningful improvements every few months. Anthropic says Sonnet 5.5 runs more than 30% faster than Sonnet 5 and costs up to 30% less for most work, at the same per-token price.

    A system built to take these upgrades gets cheaper and better over time. Taking one is routine work, but it has to be done carefully: run the new model against your own past cases, compare accuracy, cost, and speed, and switch when it wins.

    That testing and upgrading is a core part of what we do in Operate, so the improvements the labs ship reach your system without your team having to track every release.

    Find your list

    You can find your own candidates in an afternoon:

  • Write down the work your team samples, skips, or saves for month-end because doing it for every record would take too long.
  • Note the monthly volume and what a miss costs. A chargeback, a late fee, an unbilled hour, a client who waited too long for an answer.
  • Mark the items whose inputs already live in systems you control. Email, your ERP or TMS, a shared drive, the CRM.
  • Pick one and put it into production with a named person who owns the exceptions.
  • The first one shows you what the models can do on your data, and the rest of the list gets easier to prioritize from there.

    Have a list of work that was never worth doing by hand? Book a 30-minute call and we'll go through it with you: which items today's models handle well, what the monthly run cost would look like, and which one to put into production first. If it makes sense to work together, you'll get a fixed-fee proposal within a week.