Reduce coding-agent costs with OpenCode and Parasail

Alice Moore
Engineering October 2, 2026 12 min read

I use Codex and Claude Code a lot. When an agent is working in my codebase, I care much more about getting a dependable fix than saving a few cents. Paying for a capable frontier model is an easy default, but it gets expensive when a whole team is paying API rates.

So I gave the same real page-creation bug to GLM-5.3 Flash and four proprietary models: Sonnet 5, Opus 5.5, GPT-6 Sol, and GPT-6 Astra. GLM’s fix passed the browser checks for an estimated coding cost of $0.46. The proprietary models reached the same result for $2.09 to $11.89, including follow-up. That’s a big price difference for the same job.

I didn’t want to choose a team default from one bug, so I tried five more. Every workflow needed steering to finish the job, whatever it cost. If you’re going to steer anyway, invest in the setup your team steers with, and pay less for the model underneath it.

Parasail serves open models, so you can try them without running your own inference infrastructure. OpenCode lets you shape the agent’s tools and instructions around your repo and share that setup with your team. Let’s connect them, fix a bug, and see what it takes to get a result you’d actually ship.

Know when to try an open model

Open models won’t replace every frontier call, but they’re worth testing once the bill starts to shape how your team uses agents. I’d start when a few things are true:

  • Your spending grows with usage. More developers, longer sessions, and parallel agents make inference a real expense. If a subscription already covers your work, the savings matter less.
  • You can name work you do repeatedly. Bug fixes, test updates, and small refactors give you concrete assignments to compare.
  • You know how to check the result. Reproductions, tests, and browser checks let you judge whether a cheaper workflow actually finished the job.
  • You can share what works. Someone can maintain the model configuration and instructions so every developer doesn’t have to evaluate each release.

Pick candidates from the benchmarks

To build a shortlist, I’d look at how models handle coding tasks alongside what they cost to use. Artificial Analysis publishes results for Terminal-Bench 4.0, which tests agents working through terminal tasks, and SciCode, which tests scientific programming.

Here are its published model results and listed API prices, checked September 25, 2026. Rows are sorted cheapest first, with USD prices for uncached tokens.

Artificial Analysis benchmark results and listed API prices, checked September 25, 2026. USD per 1M uncached tokens, input / output.
Artificial Analysis benchmark results and listed API prices, checked September 25, 2026. USD per 1M uncached tokens, input / output.

Model

Terminal-Bench 4.0

SciCode

API price per 1M tokens, input / output

GPT-6 Luna, max effort

13%

55%

$0.10 / $0.50

GLM-5.3 Flash, effort not disclosed

33%

52%

$0.15 / $0.50

GLM-5.3, max effort

42%

59%

$1.40 / $4.40

GPT-6 Sol, max effort

44%

58%

$2.00 / $10.00

Claude Sonnet 5, high effort

5%

54%

$2.00 / $10.00

Kimi K3, max effort

13%

59%

$3.00 / $15.00

Claude Opus 5.5, max effort, default fallback

60%

67%

$4.00 / $20.00

GPT-6 Astra, max effort

59%

56%

$10.00 / $50.00

GLM-5.3 gets close to Sol on these tests at lower token prices. Flash takes the price lower again, with a stronger Terminal-Bench result than Sonnet and a similar SciCode score. Both are worth trying on real coding work, and I’ll use GLM-5.3 Flash for the walkthrough below.

Pick two or three candidates, give them work your team recognizes, and compare what it takes to finish the job. Our Kimi K3 vs. GPT-5.6 Sol comparison reaches the same conclusion across more workloads.

Connect OpenCode to Parasail

Our catalog includes models from GLM, DeepSeek, Kimi, and MiniMax behind an OpenAI-compatible API. By “open,” I mean the weights are available. Check each model’s license before you use it.

Architecture diagram: an engineer assigns and reviews work in OpenCode, the coding harness. OpenCode sends inference requests to an open model hosted on Parasail and receives responses and tool calls. It reads, edits and runs the repository and tests, and uses configured tools for browser verification.

Install OpenCode V2 2.0.15, the version I tested:

npm install -g @opencode/cli@2.0.15
opencode --version

Create a Parasail API key, then make it available to OpenCode. This Bash prompt keeps it out of your terminal history and display:

read -rs -p "Parasail API key: " PARASAIL_API_KEY
export PARASAIL_API_KEY
printf '\n'

In your repository root, create opencode.jsonc , or merge these settings into your existing configuration:

{
  "$schema": "https://opencode.ai/config.json",
  "model": "parasail/parasail-glm-53-flash",
  "providers": {
    "parasail": {
      "name": "Parasail",
      "env": ["PARASAIL_API_KEY"],
      "package": "@opencode/ai/providers/openai-compatible",
      "settings": {
        "baseURL": "https://api.parasail.io/v1"
      },
      "models": {
        "parasail-glm-53-flash": {
          "name": "GLM 5.3 Flash"
        }
      }
    }
  }
}

This uses OpenCode V2’s custom provider format and Parasail’s serverless endpoint. The key stays in your environment.

Before assigning a bug, check that the connection can handle a tool exchange:

printf 'Parasail provider check passed.\n' > provider-check.txt

opencode run --standalone \
  --model parasail/parasail-glm-53-flash \
  "Use the read tool to read provider-check.txt. Return its exact contents."

You should see a read-tool call and the file’s contents. --standalone starts a server for this run, inheriting the environment you just configured. See the CLI documentation if you normally connect to a background server.

Give both workflows the same bug

Pick a recently resolved bug for your first comparison. You already know how to reproduce it and judge a fix. Use an isolated development environment with disposable data.

Here’s the bug behind that $0.46 fix. In an open-source Notion alternative I’ve been building, clicking New page sometimes showed “Something went wrong.” Refreshing revealed that the page had been created after all.

I took the source tree from before the original fix and gave each workflow its own copy. Here’s the equivalent setup for your own pre-fix commit:

BASE=your_pre_fix_commit
mkdir -p ../coding-eval/frontier ../coding-eval/open

git archive "$BASE" | tar -x -C ../coding-eval/frontier
git archive "$BASE" | tar -x -C ../coding-eval/open

for directory in ../coding-eval/frontier ../coding-eval/open; do
  git -C "$directory" init
  git -C "$directory" add .
  git -C "$directory" -c user.name="Evaluation baseline" \
    -c user.email="eval@example.invalid" commit -m "Evaluation baseline"
done

Each copy now has a baseline you can diff against, and no later history that gives away the fix. Install dependencies and reproduce the bug in both. Keep the known fix outside the agent’s workspace and tell it not to look up the original PR.

Write down what a working fix needs to do before either run. For this pilot, the page had to open without a false error while creation was delayed, and an edited title had to survive a fresh load.

Here’s the shared assignment:

A user reports that clicking the + control beside a workspace in Content sometimes shows “Something went wrong.” Refreshing reveals that the new page was created after all.

Please investigate and fix the false failure so the user can create a page without seeing an error for a successful creation. Add a focused regression test. Validate the relevant test and, if the local development environment permits, exercise the page-creation flow.

For OpenCode, I kept the model and standing instructions in a named agent, which your team can reuse later. Save it at .opencode/agents/content-bugfix.md inside the snapshot:

---
description: Investigates and verifies bug fixes in the Content application
mode: primary
model: parasail/parasail-glm-53-flash
---

Follow the repository's Content development guidance. Reproduce the reported
behavior, make the smallest coherent fix, add a regression test, and verify
the affected user flow. Report observed evidence and any checks you could
not run.

Put the provider configuration in that snapshot too. Save the assignment as issue.md there, then run this command from that snapshot’s root:

opencode run --standalone \
  --agent content-bugfix \
  --format json \
  "$(cat issue.md)" > open-run.jsonl

Keep permission checks on, and note any help you give, like a continuation or correction.

Give your usual coding agent the same assignment in the other snapshot. Save its transcript and usage, including the model, effort setting, and tools.

Evaluation flow: both workflows start from the same unfixed source tree, assignment and acceptance checks. One runs OpenCode with Parasail, the other your usual workflow. Both go through the same tests, browser checks and diff review, then are compared on behavior, cost of every attempt, time and intervention.

See which fixes hold up, then count the cost

GLM-5.3 Flash found a race. The interface created an optimistic page while the server was still saving it. Another component immediately tried to fetch that page’s preview draft. The page didn’t exist on the server yet, so the lookup failed and the UI displayed an error.

The create request could still succeed, which is why refreshing showed the page.

The patch delayed the draft lookup until creation finished:

const creationPending = isDocumentCreationPending(document);
const drafts = usePreviewDocumentDraft(
  creationPending ? null : document.id
);

Passing null pauses the lookup, so the loading skeleton stays up until the page exists. Real lookup errors still show afterward.

With the create request delayed by 800 ms, I created five pages, watched for the false error, then edited a title and loaded the page again. GLM’s patch passed. Opus and Astra also found the race on their first coding attempts and passed those checks.

Sonnet and Sol first went after plausible backend failures after the page had been saved. Each added a passing test, but the browser still showed the error in four of five trials. After I showed them the failure and network responses, both fixed the premature draft lookup.

The estimated coding cost, including follow-up, was $0.46 for GLM-5.3 Flash, $2.09 for Sol, $3.23 for Opus, $4.49 for Sonnet, and $11.89 for Astra. All five passed the same checks. GLM needed a short permission continuation; Sol and Sonnet needed a second coding attempt. The run details include timings and the accounting for each workflow.

When you run this yourself, count every attempt, including failures and repairs, and track your own review time too. My figures cover only the API cost of the coding sessions. If your team pays for subscriptions, compare against what you actually pay.

Try to break the result

One cheap fix could be luck, so I gave seven workflows five more bugs, twice each, and mostly kept my hands off. The bugs covered editor shortcuts, large Markdown pastes, contributor permissions, MCP schema discovery, and search ranking. If an attempt fell short, it got one round of prepared feedback and a short window to repair.

None of them finished the job reliably on their own, including the expensive ones.

Some patches fixed the reported bug but broke something nearby or fell short elsewhere; the table counts those as “main fix observed.” Here’s how far each workflow got:

How far each workflow got on five more bugs, two attempts each, with coding cost in USD for all 10 attempts
How far each workflow got on five more bugs, two attempts each, with coding cost in USD for all 10 attempts

Workflow

Fully accepted, out of 10

Accepted or main fix observed, out of 10

Coding cost for all 10 attempts

GLM-5.3 Flash · OpenCode/Parasail

2

5

$4.13

DeepSeek V4.1 Flash · OpenCode/Parasail

2

4

$3.30

MiniMax M3 · OpenCode/Parasail

1

4

$9.27

Kimi K3 · OpenCode/Parasail

2

5

$52.76

GPT-6 Luna · OpenCode

2

4

$0.48

GPT-6 Sol · OpenCode

3

5

$12.56

Claude Sonnet 5 · Claude Code

0

4

$61.41

The methods companion has the full grading, costs, and run details.

The cheapest workflow wasn’t an open model. GPT-6 Luna matched GLM Flash’s two accepted fixes for $0.48, despite a much lower Terminal-Bench score, while Kimi K3 spent $52.76 on the same two. Luna, DeepSeek, and GLM Flash each cost under $5 for 10 attempts, so the bigger decision is whether to leave the frontier default at all. GLM Flash and DeepSeek earned a place in my next round of coding work.

The large-paste bug taught me the most about checks. A long Markdown draft could freeze Content before saving, and several fixes handled that big sample but broke ordinary lists or code fences. Only a smaller paste caught the damage. One of GLM Flash’s attempts solved it without feedback.

The page-creation win needed another check too. On a different baseline with broader checks, only two of 14 attempts passed, both proprietary; none of the eight open-model attempts did.

What helped was steering. Of the 12 accepted fixes, 11 came after the feedback round. You pay that cost with any model, so pay it on the cheap one, and write it down so your team doesn’t pay it twice.

Turn your steering into a team setup

Save the steering that turned those runs around. OpenCode lets you keep it in the repo next to the model choice, so your teammates start where you left off.

If you enjoy keeping up with model releases, great. Your whole team shouldn’t have to.

You can share the verification procedure too. For the page-creation example, create .opencode/skills/verify-page-creation/SKILL.md :

---
name: verify-page-creation
description: Verify changes to page creation and optimistic page loading
---

1. Read the repository's development and test instructions.
2. Reproduce the reported failure on the baseline.
3. Run the focused tests for page creation and draft recovery.
4. In the browser, delay a successful create request by 800 ms.
5. Repeat creation five times. Check for a false error while it is pending.
6. Edit a created page, reload it, and confirm the edit persisted.
7. Report commands, observed results, and any verification you could not run.

Add the repo’s exact commands or browser script alongside it. In the pilot, I handled step 4 with Playwright:

await page.route(/actions\/create-document/, async (route) => {
  await new Promise((resolve) => setTimeout(resolve, 800));
  await route.continue();
});

This writes down the checks from the pilot so the agent can run them itself; the original runs didn’t have it. Add to it as you hit new page-creation bugs.

OpenCode discovers repository skills. Add this instruction to the named agent: “Use the verify-page-creation skill when a change affects page creation or optimistic page loading.”

Now the next developer has both a model configuration and a concrete way to check its work. When a new model becomes worth trying, update its provider entry and the named agent’s model, then rerun the saved tasks against the same checks.

Choose who runs the model

A proprietary model comes from its lab or one of a few cloud platforms. An open model can come from almost anyone with GPUs. When I checked on September 28, OpenRouter listed 33 endpoints for GLM-5.3 Flash from 31 providers. They ran it at 4-bit or 8-bit precision, or didn’t say which, at a range of prices. Our guide to open-weight inference providers compares their pricing, commitments, regions, and precision.

That’s nice while you’re shopping around, but a team default should run on the deployment you tested. Every open-model run in this article used Parasail’s API directly. If you switch providers, rerun your saved tasks before you assume the results carry over.

Routers such as OpenRouter make shopping easy: one API key covers many models and providers. By default, OpenRouter picks a provider for each request, favoring cheaper ones, and falls back to another provider if the first is unavailable. You can restrict requests to one provider, such as Parasail, and turn off fallbacks. OpenRouter also has controls for data retention and regional routing.

Once a model becomes your team’s default, or part of internal tools other employees rely on, I’d switch to the provider’s own API. Every request reaches the deployment you evaluated, and your prompts go to one company instead of two. At Parasail, we don’t train on your inputs or outputs, and we keep personal data from your requests only as long as we need to respond. Going direct also lets you run private Hugging Face models, so you can try your own supported fine-tuned weights behind the same agent and checks.

Before you move a team off Claude Code or Codex, compare your provider’s limits with your team’s peak usage, since parallel agents add up. Our standard serverless accounts allow 500 requests per minute, and Enterprise accounts have no request limit. Serverless runs on shared capacity, so latency can vary. With an Elastic Endpoint or a dedicated deployment, you reserve capacity: your limit is what you provision, and latency is more consistent. Our guide to choosing an inference architecture covers when each option makes sense.

Your security team will also ask where your prompts are processed and which compliance reports cover the provider. Parasail is SOC 2 Type II compliant. All four open models I tested come from labs based in China, but open weights run wherever the provider runs them, and Parasail offers US-only serving for customers who need it.

Swap models, tools, and deployments as you learn

This setup lets your team change almost every piece, starting with the model underneath. That flexibility is also a hedge against depending on a single frontier API.

Start with the serverless API from this walkthrough, where trying another model family is a small config change.

OpenCode lets you mix models to fit the job. Each agent can use its own model and provider while keeping its instructions and tools, and you can switch a session’s model with /models , including to one from a different provider. Start an investigation on GLM-5.3 Flash, move to a larger model if it stalls, and bring in a frontier model only for the review.

You can also add repository-specific tools and hooks with plugins, or change the harness itself under its MIT license. If an investigation needs another machine’s workspace or compute, OpenCode can connect to a remote server.

Whichever model you choose, you’ll still be steering it. Write that steering down once, share it with your team, and run it on a model that costs a fraction of the frontier default. Pick one recurring task, try an open model on Parasail, and check the result. Once it does the job reliably, make it the easy choice. Frontier pricing should earn its place in the work that’s left.

Related blog posts

Blog

Start building today

Instantly run any open model — popular or specialized.