Agent costs from real tasks
Explore the task, the models used and the recorded cost together.
These are selected examples, not the limits of what agents can do. Agent capabilities extend beyond the tasks, industries and regions shown here, and can be adapted to different workflows and business needs.
Agent costs vary dynamically with the scope of the request, conversation context, models and tools used, response length and usage volume. The amounts below are reference examples from verified historical execution records. External service and infrastructure fees should be considered separately.
How to read the costs
For full conversations, Calculated total sums the verified execution records for every displayed turn, including continuation and delegated work. No turns are removed by percentile filtering, and delegated costs are counted only once.
Amounts are converted to USD from agent execution records matched to the request. Models were verified from the same execution traces. All models are listed when a task used more than one. Any continuation after approval is included in the cost.
For the original single-task selection, for agents with at least 10 positive-cost records, the lowest and highest 10% were excluded from selection, rounding the number of records down. No percentile filter was applied to agents with fewer records. Each selected task retains its recorded cost; it has not been replaced by an average. Different models may have different prices.
These amounts are neither fixed prices for future usage nor a complete service price. External services such as image generation, hosting and other infrastructure fees should be considered separately. The examples come from different dates; each detail page includes the record date and cost scope.
Client, brand and personal information has been removed. Examples include summaries and complete available conversations. Full conversations preserve every user and agent message in order, translated into English where needed. The status reflects the outcome reported at the time.
| No. | Industry / region | Agent | Task summary | Models used | Details | |
|---|---|---|---|---|---|---|
| 1 | A coffee & hospitality chain in MENA | Operations assistantSales analytics | Store comparison Show the five stores with the highest completed-order revenue over the last 90 days, including average order value. Ranked five stores by revenue, order count and average order value. Store A led with AED 8,429 from 73 orders and an average order value of AED 115.47. Store B generated AED 5,088 from 85 orders. Converted the recorded minor currency units to AED. Analysis completed | Claude Sonnet 5 | $0.0598261 run · 2 model calls | View task → |
| 2 | A coffee & hospitality chain in MENA | Operations assistantCampaign management | Campaign status Show the live campaign and the status of campaign drafts. Distinguished a campaign published on web and mobile from a second campaign still in draft. Reported one conversion and AED 28.80 in attributed value for the published campaign. Found the customer-facing page without changing either campaign. Status reported | Claude Sonnet 5 | $0.0386031 run · 2 model calls | View task → |
| 3 | A coffee & hospitality chain in MENA | Operations assistantCampaigns and content | Campaign image generation Generate a 3:2 hero image for an existing coffee campaign, with no text or logos, and attach it to the draft. Generated a 1536×1024 hero image featuring a coffee bag and a ceramic cup on warm beige stone. Attached the image to the campaign’s draft content and prepared alternative text. Reported completing the image task while keeping the campaign unpublished. Image attached · campaign in draft | Claude Sonnet 5 | $0.0472861 run · 2 model calls | View task → |
| 4 | A coffee & hospitality chain in MENA | Operations assistantOrder operations | Order operations Show the status and total of the three most recent active orders. Listed three orders, all with received status, totaling AED 28.80, AED 20 and AED 67 respectively. Included the store, creation time and pickup or table-service details for each order, helping the operations team identify orders awaiting their next step. Analysis completed | Claude Sonnet 5 | $0.0347891 run · 2 model calls | View task → |
| 5 | A coffee chain in Türkiye | Sales analytics assistantSales analytics | Branch performance analysis Find underperforming coffee-chain branches over the last 28 days and explain the method. Compared the last 28 days with the preceding period, assessing revenue against the network median and growth against the network average. Weighted the median gap at 60% and the growth gap at 40%. Explained the prioritization using Store A, where orders fell from 107 to 30 and revenue from TRY 32,385 to TRY 8,735. Analysis completed | Claude Haiku 4.5 + Claude Sonnet 5 | $0.1149511 run · 4 model calls | View task → |
| 6 | A coffee chain in Türkiye | Loyalty assistantLoyalty and segmentation | Loyalty segment preview Preview Gold members with at least 10 orders for a loyal-customer segment, then request approval to save it. Found 29 customers matching Gold membership and at least 10 orders. Reported TRY 402,143 in total spending and TRY 13,867 per customer, along with the order-count range. Requested approval to save the segment; the response did not claim that it had already been saved. Awaiting approval to save | Claude Haiku 4.5 + Claude Sonnet 5 | $0.0995891 run · 4 model calls | View task → |
| 7 | A restaurant project in North America | Menu advisorProduct advice | Product navigation Give me a direct link to the birria tacos product page. Located the three-piece tacos item on the menu and directed the user to its product page at a price of USD 14.50. The client-specific website link has been removed from this shared version. Product and page found | Claude Sonnet 5 | $0.0187541 run · 2 model calls | View task → |
| 8 | A restaurant project in North America | Menu advisorProduct recommendations | Vegetarian menu recommendations Recommend two vegetarian dishes with prices. Recommended a tomato-and-cheese pizza at USD 14.99 and spicy miso vegetable ramen at USD 15.99. Explained the pizza’s tomato, mozzarella and basil ingredients and the ramen’s tofu, mushrooms, corn and chili oil. Presented both options with dietary preference, ingredients and price. Recommendations provided | Claude Sonnet 5 | $0.0249081 run · 2 model calls | View task → |
| 9 | A restaurant project in North America | Menu advisorLoyalty management | Loyalty program How do reward points and membership tiers work? Explained that customers earn one point for each USD 1 spent. Listed four membership thresholds at 0, 500, 1,200 and 2,500 points, mapping benefits such as free delivery, bundle discounts and double points on selected days to their respective tiers. Rules explained | Claude Sonnet 5 | $0.0219051 run · 2 model calls | View task → |
| 10 | A restaurant project in North America | Menu advisorProduct recommendations | Budget-based menu planning What can two people eat for less than USD 30 in total? Recommended a sharing platter with two drinks for USD 26.97, or birria tacos with chili cheese fries for USD 25.49. Showed individual prices and combined totals to create shareable combinations within the budget for two people. Budget options provided | Claude Sonnet 5 | $0.0696601 run · 2 model calls | View task → |
| 11 | A restaurant project in North America | Menu advisorProduct advice | Product and price information List two menu items with their prices and ingredients. Described a double-patty burger priced at USD 15.99, including fries and its main ingredients. Offered ramen at USD 16.50 as a second option, explaining its broth, meat and egg components. Supported product choice with both ingredient and price information. Product information provided | Claude Sonnet 5 | $0.0655911 run · 2 model calls | View task → |
| 12 | A multi-brand retail project in Türkiye | Risk analystRisk analysis | Promotion abuse analysis Explain the risk score and evidence for a suspicious promotional order. Reported a score of 73/100 and the system’s medium-risk classification for an order in a test scenario. Cited triple the usual points earning, a TRY 45,490 basket and device/IP indicators. Left the approval or rejection decision to a human moderator rather than automatically rejecting the order. Awaiting human reviewTest scenario | Claude Haiku 4.5 | $0.0265801 run · 2 model calls | View task → |
| 13 | A fashion e-commerce project | E-commerce analystSales analytics | Weekly revenue distribution How did revenue change over the last seven days? Identified 13 paid orders during the seven-day period. Showed that nine orders generated USD 2,945 on July 29, while three days had no recorded orders. Explained the uneven daily distribution without presenting a single-day spike as sustained growth. Trend analyzed | Claude Opus 4.8 | $0.0753751 run · 3 model calls | View task → |
| 14 | A quick-service restaurant chain in MENA | Menu advisorProduct advice | Menu options within a budget What would you recommend for two people for less than AED 80? Recommended a chicken-and-pasta meal with two drinks for AED 60, two sandwich meals for AED 50, or a six-piece chicken sharing meal with a one-liter drink for AED 70. Presented the contents and combined price of each option within the budget. Budget options provided | Claude Sonnet 5 | $0.0625751 run · 2 model calls | View task → |
| 15 | A coffee chain in Europe | Store analystStore operations | Store status How many stores are currently open? Reported that five of six stores were marked open and one was marked closed. Explained that the result came from the open/closed field in store records, and suggested checking store status and opening hours if a closure was unexpected. Stores checked | Claude Sonnet 5 | $0.0668151 run · 3 model calls | View task → |
| 16 | A coffee chain in Europe | Promotions assistantCampaigns and content | Promotion rules Show all active promotions. Explained the scope and conditions of four promotions. Identified Friday Points Benefit as a display badge only, and the EUR 1 discount on orders above EUR 10 as limited to five uses per customer. Also separated the segment, product and usage conditions of targeted cold-coffee and milk-coffee discounts. Analysis completed | Claude Sonnet 5 | $0.2033661 run · 3 model calls | View task → |
| 17 | A coffee chain in Europe | Store operations assistantStore operations | Store change with approval Close the selected store. Matched the store and submitted the closure for approval. After approval in the same conversation, reported completing the change and verifying that the store was marked closed. The cost includes both the initial request and the continuation after approval. Approved · closure verified | Claude Sonnet 5 | $0.0597892 runs · 5 model calls | View task → |
| 18 | A coffee chain in Europe | Store operations assistantCampaigns & content | Sales analysis, banner and daily operations Find the least-selling products and create a red banner; then review the product ranking, active promotions and daily brief. The assistant publishes a banner, lists 10 products and 2 active promotions, then saves a daily brief that flags a data gap. Banner and brief reported saved; the later brief flags a data sync concern. | Claude Sonnet 5 | $0.443081Calculated total4 runs · 12 model calls | View full conversation → |
| 19 | A retail customer-data project in Türkiye | Data quality assistantData quality | Name, email and phone conflicts Investigate name conflicts, provide email–phone examples and profile customers with conflicting phone records. The assistant examines all three cases, reports affected populations, supplies anonymized examples and directs record corrections to the case workflow. Analysis completed; no customer records modified. | DeepSeek V4 Flash | $0.006173Calculated total3 runs · 6 model calls | View full conversation → |
| 20 | A quick-service restaurant chain in MENA | Operations assistantCampaign management | Campaign opportunity and personal winback offer Identify an untargeted customer group and prepare a 20% next-order winback offer. The assistant identifies 625 customers with no orders and 5 dormant customers, checks the chosen cohort and asks which accounts should receive personal coupons. Awaiting audience confirmation; no coupon creation reported. | Claude Opus 4.8 + Claude Sonnet 5 | $0.514483Calculated total5 runs · 15 model calls | View full conversation → |
| 21 | A fashion e-commerce project | Growth analystCampaigns & content | From slow-selling products to a flash sale and win-back audience Analyze weak sellers, scope a clearance campaign, launch a 10-product flash sale and create a marketing-consented win-back audience. The conversation covers sales and stock analysis, several offer revisions, campaign delegation, website placements and a 1,812-member win-back segment. Campaign reported live; email sending and price restoration still require follow-up. Win-back segment created. | Claude Opus 4.8 | $3.393595Calculated total19 runs · 62 model calls | View full conversation → |
| 22 | A coffee chain in Türkiye | Sales and retention analystSales analysis | Store decline diagnosis and retention campaign preparation Find underperforming stores, then prepare a win-back campaign for their inactive customers. The assistant compares two 28-day periods, diagnoses falling order volume, identifies 119 inactive customers and prepares a 50-point reward campaign for approval. Campaign and audience submitted for human approval; publication not completed. | Claude Haiku 4.5 + Claude Sonnet 5 | $0.567211Calculated total2 runs · 16 model calls | View full conversation → |
| 23 | A coffee chain in Türkiye | Loyalty assistantLoyalty & segmentation | Gold audience preview and approval workflow Preview Gold members with at least 10 orders and request approval to save the audience. An approval test previews 29 customers and their spending, then creates the segment approval card across two execution turns. Test scenario: awaiting approval; permanent segment creation not confirmed. | Claude Haiku 4.5 + Claude Sonnet 5 | $0.121518Calculated total2 runs · 6 model calls | View full conversation → |
| 24 | A coffee chain in Europe | Promotion and operations assistantCampaigns & content | Promotion draft, product ranking, banner and daily brief Prepare a category discount, review the 10 lowest-selling products, create their banner and generate a daily brief. The assistant prepares a 15% discount for approval, reports product sales, publishes a banner and flags a sales-reporting gap in the daily brief. Promotion awaits approval; banner published and daily brief saved with a data-gap note. | Claude Sonnet 5 | $0.480275Calculated total5 runs · 13 model calls | View full conversation → |
| 25 | A restaurant project in North America | Operations assistantSales analysis | Lowest-selling item analysis and campaign scoping Identify the weakest-selling item and start defining a campaign for it. After a specialist session error, the assistant queries the data directly, compares five low-selling items and asks for campaign offer, audience and timing. Analysis completed; campaign parameters await confirmation, no campaign created. | Claude Opus 4.8 + Claude Sonnet 5 | $0.291227Calculated total3 runs · 11 model calls | View full conversation → |
| 26 | A retail customer-data project in Türkiye | Customer-data analystData quality | Investigating two brand-specific name variants Inspect a customer's names across two brands and clarify the exact underlying name variants. The analyst performs 11 model calls over two turns, identifies the name-conflict case and explains that the available tools expose variant counts but not the individual raw names. Conflict confirmed; exact raw variants remain unavailable through these tools. No data modified. | DeepSeek V4 Flash | $0.004812Calculated total2 runs · 11 model calls | View full conversation → |
| 27 | A coffee & hospitality chain in MENA | Operations assistantCampaign management | Low-seller campaign draft and approval continuation Find a low-selling product and prepare a draft 10% campaign for web and mobile without publishing it. An approval verification scenario selects a zero-sales coffee product, prepares the draft campaign and explains the planned records and publication boundary. Test scenario: draft workflow prepared; no public campaign page reported. | Claude Sonnet 5 | $0.113695Calculated total2 runs · 4 model calls | View full conversation → |
| 28 | A coffee & hospitality chain in MENA | Operations assistantCampaign management | Campaign approval rejection and corrected outcome Prepare a 10% draft campaign for a weak product and route it through approval. The approval test identifies a zero-sales product, submits a campaign and then corrects the initial pending claim after the operator rejects the action. Test scenario: rejected by operator; no campaign or CMS record saved. | Claude Sonnet 5 | $0.101732Calculated total2 runs · 4 model calls | View full conversation → |
| 29 | A coffee chain in Europe | Store operations assistantStore operations | Store availability and cross-checked daily operations List open stores, then produce the daily brief using sales and order-queue data. The assistant lists 11 stores, finds one closed store and saves a brief that contrasts missing sales data with active and completed orders. Store status reported and brief saved; sales-data pipeline needs investigation. | Claude Sonnet 5 | $0.286843Calculated total2 runs · 7 model calls | View full conversation → |
| 30 | A coffee & hospitality chain in MENA | Retention operations assistantLoyalty & segmentation | Churn-risk portfolio and segment approval Identify customers at risk of churning, then request an active segment for repeat buyers inactive for at least 90 days. The assistant profiles 400 customers, identifies 119 high-risk and 39 medium-risk cases, submits a segment approval request and reports the operator's rejection. Verification scenario: analysis completed; segment rejected and not created. | Claude Sonnet 5 | $0.214262Calculated total3 runs · 5 model calls | View full conversation → |