AI cost: from token rates to accepted work
Compare Sol, Terra, Luna and Sonnet 5 using the right units. Include output, retries, tools and review before discussing returns.

The lowest rate per million tokens does not, by itself, identify the most economical option. A task may require retries, produce long answers or use additional tools. The useful decision metric is the expense of obtaining an accepted result at the required quality.
Official tables checked on September 21, 2026 listed these standard text rates per million uncached input and output tokens: GPT-5.6 Sol: US$4 / US$20; Terra: US$2 / US$12; Luna: US$0.20 / US$1.20. For Claude Sonnet 5: US$2 / US$10. These are standard API references; other processing modes, long context, caching and tools have their own conditions.
Calculate with the correct units
For uncached text, multiply input tokens by the input rate and output tokens by the output rate, dividing each token quantity by one million. Then add the two amounts. Do not apply the input rate to the entire volume, because output can carry a different price.
As an arithmetic example, not measured project usage, consider a Sol request with 100,000 input tokens and 20,000 output tokens. At the rates above, the two amounts are US$0.40 and US$0.40: US$0.80 before any other applicable charges. This does not predict how many tokens a real task will consume or how many attempts will be needed.

Include attempts that did not become a delivery
If three responses were rejected before one solution was accepted, all belong to the cost of that task. Separate the reasons: missing context, ambiguous instructions, tool errors or an unsuitable solution. This classification helps identify where the process needs improvement. A cheaper model does not solve every one of those problems.
Record unfinished work too. Dividing a bill only by successful tasks while omitting abandoned work can hide waste. A useful view includes tasks started, accepted and interrupted, with their associated expense. Keep the definition of acceptance stable throughout the comparison, rather than lowering it when a result is inexpensive.
Include tools and human effort without inventing savings
Search, execution, storage, subscriptions and infrastructure may add charges depending on the product. Consult the relevant billing terms and avoid confusing API rates with subscription limits. Two services can use the same model while charging for the overall experience differently.
Review time matters as well. Record observed minutes and the intervention required. If converting time into money, state the rate and whether it represents an actual cost or a planning assumption. Apparently freed hours do not automatically become reduced payroll or additional revenue. Someone must decide how that capacity can be used.
Examine spending by task type
Group similar work before comparing averages. A short classification and a lengthy code investigation should not share a benchmark without context. Begin with categories showing high expenditure or repeated failures, while retaining quality samples so lower spending does not conceal worse outputs.
The guide to Sol, Terra and Luna organizes alternatives. The development comparison with Claude Fable 5 helps define accepted work. Returns become meaningful only when costs and benefits have an identifiable basis, rather than relying on a general promise of productivity.
About the author
Tiago F SantiagoComments
No comments yet
Share a question or an experience related to the article.


