In our eight-case Task Automation Agent relevance benchmark, outputs that referenced the described process domain increased from 1 of 8 cases (13%) before the prototype change to 8 of 8 (100%) after it. The eight cases covered eight different process domains, so the benchmark produced no inventory-specific pass/fail rate. This measured domain relevance only—not whether any generated automation plan was good, complete, safe, or ready to run.
Direct answer
Our measured result was a change from 1 relevant output in 8 cases before the prototype change to 8 relevant outputs in 8 cases after it.
The benchmark tested whether an output referred to the process domain described in the task. It did not assess whether the resulting automation plan would operate correctly in inventory, order, fulfillment, or finance processes.
What we tested in the task-relevance benchmark
We captured this relevance benchmark for our Task Automation Agent on 2026-08-27.
The sample consisted of eight recurring-task descriptions written by us across eight process domains. Inventory and order automation were part of the benchmark context because recurring operational tasks often begin with a plain-language description: review an exception, validate a record, route a request, or take a follow-up action.
Earlier generator behavior failed a basic test before any plan could be evaluated: it did not reliably refer to the process domain the user had actually described.
So we tested the first gate:
- Provide a recurring-task description.
- Inspect the returned task content.
- Check whether the output references the process domain in that description.
That is the full claim behind this benchmark. It is not a claim that there is one right automation plan for each task, because there is not.
The benchmark measured domain relevance, not plan quality
A relevant reference is a low bar, but it is still a necessary one.
For an inventory-related recurring task, a response should refer to inventory rather than return generic record-retrieval language. For an order-related task, it should refer to the order process rather than reuse the same generic task text. That does not tell us whether the logic handles reservations, available-to-promise rules, allocation conflicts, cancellations, partial shipments, substitutions, or returns.
It also does not establish whether a plan is efficient, safe, governed appropriately, or suitable for execution.
This distinction matters because inventory and order workflows depend on operational facts outside a task description. For example, Shopify assigns orders to fulfillment locations through configured routing rules and available inventory across channels, while Salesforce Distributed Order Management evaluates fulfillment locations and reserves inventory for order items. Those are operational decisions with dependencies beyond a domain reference. See Shopify’s fulfillment-location documentation and Salesforce’s distributed order-management documentation.
Why inventory and order workflows were evaluated as recurring task descriptions
Inventory and order work is often expressed as recurring work rather than as a technical specification. A team may describe a repeated task in operational language and expect the output to reflect that context.
That expectation is reasonable, but our earlier prototype behavior exposed why it must be tested. If a system returns the same check, condition, and action regardless of whether the description concerns inventory, orders, or another domain, the output is not yet responsive enough to assess as an operational plan.
This is also why we keep the Task Automation Agent separate from Intelligent Document Processing. Extracting information from a document and generating a task response are different steps. A relevant task reference does not demonstrate reliable extraction from emails, PDFs, EDI messages, marketplaces, or other order inputs.
The measured result: 1 of 8 correct before, 8 of 8 after
The benchmark contained eight cases. These are the only measurements reported from the captured test.
| Measure | Before prototype change | After prototype change |
|---|---|---|
| Sample size | 8 cases | 8 cases |
| Outputs that passed the process-domain relevance check | 1 | 8 |
| Relevance accuracy | 13% | 100% |
The figures describe this eight-case sample only. They should not be projected to other tasks, customers, systems, or automation workflows.
Before: only one of eight outputs referenced the described domain
Before the prototype change, 1 of 8 outputs passed the relevance check.
The problem was not a subtle disagreement over workflow design. The returned content was not being driven by the described task domain. As a result, most outputs failed before there was a meaningful basis to discuss whether a proposed process was operationally sound.
A schedule may have looked different from one returned task to another, but the task-specific substance did not.
After: all eight outputs passed the relevance check
After the prototype change, 8 of 8 outputs passed the relevance check in this benchmark.
That means every tested output referenced the process domain actually described in its recurring-task prompt. It does not mean each output contained a usable automation sequence. We did not score completeness, exception handling, data validation, approval logic, routing decisions, or run readiness.
The result is narrowly useful: it shows that, in these eight cases, the prototype cleared the relevance gate that its earlier behavior had missed.
What surprised us about the failed outputs
The failure mechanism was direct: the check, condition, and action lines were hardcoded strings.
Every task returned:
Retrieve open records
That happened regardless of the task description.
The implementation produced the appearance of a task response, but its central content was fixed. The description supplied by the user did not control the substantive task lines.
The task description did not control the check, condition, or action
The observed hardcoding affected the parts that should have made the output task-specific:
- the check;
- the condition; and
- the action.
Each returned the same hardcoded content: “Retrieve open records.”
This is exactly the type of failure a relevance check is intended to expose. An inventory description can produce an inventory-sounding schedule and still fail if the actual check, condition, and action remain generic. An order description has the same problem if the output does not change beyond timing.
The benchmark did not assess whether any revised task logic was operationally adequate. It assessed whether the returned output finally referenced the described domain.
Schedule variation did not make the returned task relevant
Only the schedule varied in the failed outputs.
That variation did not fix the hardcoded check, condition, or action lines. A schedule answers when a task might run; it does not establish what the task evaluates or what it does.
For inventory and order work, this distinction is especially important. A recurring schedule cannot compensate for content that ignores whether the described task concerns stock, orders, routing, customer enquiries, or another process domain. For order-status questions, the relevant process context may be closer to the workflows shown in our Order & Customer Enquiry Agent than to a generic record-retrieval instruction.
What the result does—and does not—prove
This result supports one conclusion: process-domain relevance improved within our captured eight-case benchmark.
It does not establish that generated automation plans are correct, complete, efficient, safe, or ready to run. There was no single correct plan for each description, and plan quality was not measured.
It also does not establish inventory accuracy, order accuracy, oversell prevention, service-level performance, cost outcomes, time savings, or any other business result.
A relevant reference is not the same as a working inventory workflow
An output can mention inventory and still be operationally weak.
It might treat on-hand stock as available-to-promise stock without accounting for reservations, safety stock, damaged goods, allocations, or in-transit stock. It might choose a nearest-location routing rule while overlooking split shipments, service requirements, carrier constraints, margin, stock aging, or customer commitments. It might cover order intake but omit cancellations, returns, backorders, substitutions, and inventory adjustments.
Those are not defects this benchmark measured. They are examples of why a relevance result must not be treated as operational approval.
Data quality is another separate concern. Automation acting on uncorrected inventory data can propagate errors rather than resolve them. That issue requires its own validation, beyond whether a generated task refers to inventory.
Why this evidence cannot establish broader automation performance
Eight cases are not evidence of broad automation performance. They are eight descriptions written by us, captured at one point in time, and assessed against one criterion: whether the output referenced the process domain described.
The benchmark did not use a control group, test messy source documents, score exception handling, or measure post-implementation outcomes. It did not compare deterministic execution rules with AI-assisted recommendations, and it did not test confidence thresholds, human escalation, rollback procedures, or audit trails.
For that reason, neither 13% nor 100% should be generalized beyond these eight cases.
How to use this evidence when reviewing an automation prototype
Use this result as a template for a narrow first review, not as a procurement scorecard.
Provide recurring-task descriptions that reflect the work your team actually needs reviewed. Then inspect the returned content before debating whether the proposed automation logic is sophisticated. If the same generic check, condition, and action appear across different task domains, stop there: the system has not passed the relevance gate.
After that gate, validate the operational plan separately against the data, rules, exceptions, and approvals in your environment.
Check task-specific content instead of schedule changes alone
For each recurring-task description, inspect whether the following content changes in response to the described domain:
- Check: what records, conditions, or signals are being reviewed?
- Condition: what determines whether the task advances?
- Action: what is proposed or performed after the condition is met?
Do not accept timing changes as evidence of relevance. Our failed outputs varied only in schedule while the substantive lines remained hardcoded as “Retrieve open records.”
For inventory and order scenarios, task-specific review should include whether the output distinguishes inventory truth from available inventory, and whether it names the actual order or fulfillment context described. That still does not validate the logic; it only confirms that the response is about the stated work.
Separate relevance testing from operational approval
Treat relevance testing and operational approval as separate reviews.
First, test whether the output refers to the described process domain and whether its check, condition, and action vary with the task. Our eight-case benchmark is evidence for that first review only.
Then validate the plan against master data, including SKU, unit of measure, location, customer, pricing, and available-inventory data. Review exception queues and escalation paths. Assess routing and allocation choices against the organization’s own commitments and constraints. Confirm what requires human approval and what can execute under deterministic rules.
You can review these capabilities in our live demos, but a demo does not replace task-specific validation in the systems where orders, inventory positions, and fulfillment commitments are maintained.
Frequently asked questions
What did the eight-case task-relevance test measure?
The test measured process-domain relevance in eight recurring-task descriptions written by us. We checked whether each Task Automation Agent output referenced the process domain actually described in the task. It did not measure whether an automation plan was good, complete, safe, efficient, or ready to operate in an inventory or order environment.
How many cases passed the relevance check before and after the prototype change?
Before the prototype change, 1 of 8 cases passed the relevance check, reported as 13%. After the change, 8 of 8 cases passed, reported as 100%. These figures apply only to the captured eight-case benchmark on 2026-08-27 and should not be generalized to other tasks or deployments.
Why did the earlier automation outputs return “Retrieve open records”?
The earlier outputs returned “Retrieve open records” because the check, condition, and action lines were hardcoded strings. The task description did not control those substantive lines, so every task received the same content regardless of the process domain described. Only the schedule varied, which did not make the returned task relevant.
Does a 100% relevance result prove that the automation plans work?
No. The 100% result means all eight outputs referenced the described process domain in this benchmark. It does not establish that any plan handles inventory data, reservations, routing, exceptions, approvals, or execution safely. There was no single correct plan for each case, and the benchmark did not measure plan quality or operational outcomes.
Researched and drafted with AI assistance, reviewed and fact-checked before publication. Reviewed by M. Haroon.
Try the related live prototype