SOLUTION USE CASE

AI Usage Analysis: Large-Model Tokens and Cost in One Place

Per-platform consoles are separate and the cloud bill is just a total — no tokens, no models. Mof unifies large-model usage on GCP, Alibaba Cloud, Tencent Cloud, Volcengine, Huawei and Baidu into one table: token detail, cache hit rate and real payable by model and native dimensions.

Status quo · every AI platform on its own

Bailian / Ark / Qianfan separate

Bill total only

No tokens / models

Compute cache rate yourself

Scattered, totals only — you can’t tell which model burns the spend

With Mof · AI usage in one table

With Mof · AI usage in one table

By model
Token detail
Cache hit rate
Real payable

Read-only, multi-platform AI usage unified in one place

Why teams use Mof for AI usage analysis

Large-model usage keeps growing and spend rises fast, yet “how much did AI cost this month, on which model, for which app?” is hard to answer: Bailian, Ark, Qianfan and Vertex each sit in their own console with different schemas; you can’t see token detail or cache hit rate to optimize; and the cloud bill shows AI as a single total.

Mof is a FinOps platform for teams running more than one cloud. Its AI usage analysis connects large-model usage on GCP, Alibaba Cloud, Tencent Cloud, Volcengine, Huawei Cloud and Baidu AI Cloud read-only into one table: slice by model and each platform’s native dimensions, see input / cached / output token detail, cache hit rate and amount payable, with trends, composition and drill-down — turning AI cost from “one lump sum” into “clear, splittable and reducible.”

How teams view AI usage today — each with a catch

Per-platform consoles, the bill total, pulling APIs — each shows a bit, but none is cross-platform, token-level or cache-aware.

Per-platform consoles
  • One console per platformBailian, Ark, Qianfan, Vertex — each separate
  • Different schemasToken counting, currency, periods all differ
  • No whole pictureTotal AI spend and which models? Hard to tell
Just the cloud bill total
  • Totals onlyAI is one line, no by-model / by-app view
  • No token detailinput / cached / output not separated
  • Can’t optimizeNo idea which model or app is burning spend
Pull APIs yourself
  • Each platform differsEvery usage API integrated separately
  • Compute metrics yourselfCache hit rate, unit price and more
  • Not sustainableModels change, accounts grow — rebuild it

How Mof does it

Unify multi-platform AI usage into one table, by model and each platform’s native dimensions, with token detail, cache hit rate and real payable.

1Connect AI cloud accounts
2Normalize usage & cost
3Slice by model / dimension
4Optimize model & cache
By model + native dimensions
Slice the same AI usage by model, and keep each platform’s native dimensions (e.g. Bailian’s app space / finance unit, other clouds’ Region).
Token detail
See input / cached / output / embedding tokens by type — trends and composition — so you know exactly where tokens go.
Cache hit rate
Built-in cache hit rate and trend — higher hit rate, more saved — so you spot the headroom at a glance.
Real payable
Shown as amount payable, with multi-currency normalized at FX to your chosen currency — comparable across platforms and models.
Cross-cloud, cross-account
Large-model usage on GCP, Alibaba Cloud, Tencent Cloud, Volcengine, Huawei Cloud and Baidu AI Cloud, unified into one table.
Trends + drill-down
Usage and cost by trend and composition over time, with a detail table you can expand per model — see what rose and which model is pricey.

Platform consoles / bill vs Mof

DimensionPlatform consoles / billMof
CoverageEach platform on its ownGCP / Alibaba / Tencent / Volcengine / Huawei / Baidu in one table
GranularityBill total onlyBy model + each platform’s native dimensions
Token detailNoneinput / cached / output / embedding
Cache hit rateCompute it yourselfBuilt-in, with trend
Group / allocateNoneNative dimensions + cost allocation / virtual tags
CurrencyPer-platform currenciesFX-normalized to your chosen currency
Drill-downNoneDetail table, expand per model

AI usage analysis FAQ

Which platforms and models does AI usage analysis support?

Large-model usage on GCP, Alibaba Cloud, Tencent Cloud, Volcengine, Huawei Cloud and Baidu AI Cloud — covering common models such as Alibaba Bailian (Qwen), Volcengine Ark (Doubao), Baidu Qianfan and Gemini / Claude on GCP — synced with the bill and usage.

Can I see token-level detail?

Yes. Mof breaks tokens down by type (input, cached, output, embedding) and provides usage / cost trends and composition by model and each platform’s native dimensions (e.g. Bailian’s app space / finance unit, other clouds’ Region), plus a detail table you can expand per model.

Why does cache hit rate matter?

Cached tokens are usually much cheaper than regular input, so a higher cache hit rate means more savings. Mof shows the cache hit rate and its trend, so you can tell whether prompt / context reuse is paying off and where the headroom is.

Can I allocate AI cost by app or team?

Yes. Mof keeps each platform’s native dimensions (e.g. Bailian’s app space / finance unit, other clouds’ Region) and combines them with its own cost allocation and virtual tags to land AI spend on the right app, team or project.

How does it connect? Is it safe?

Through read-only permissions on each cloud / AI platform for usage and billing data — no write access, and nothing in your model services is changed — to unify multi-platform AI usage in one place.