How Reflow AI lowered inference costs for document ingestion and OCR by 70% with Parasail

$0

idle GPU cost

70%

lower inference spend

< 2s

end-to-end response latency

Reflow AI
“Before Parasail, we were renting GPU servers, but our load wasn't stable enough to use them efficiently. Moving to Parasail’s Dedicated Serverless helped us eliminate idle GPU costs, saving us 70% on our inference spend.”
Akbay Tabak Head of Engineering, Reflow AI

Overview

Turning how work happens into operational intelligence

Reflow AI turns computer interaction data into operational intelligence and automation opportunities for enterprise teams. Endpoint apps installed on employees' computers capture their interactions so Reflow can understand how work happens and identify where operations can improve or be automated.

As an agentic product, inference is core to their operations.

The problem

Paying for GPUs that sat idle half the week

Reflow needs to process large volumes of document and screen data through OCR, ingestion, and analysis. But inconsistent usage throughout the week meant that Reflow was paying for capacity that was often sitting idle. Their main challenge was finding a way to meet their performance requirements without paying for unused GPUs.

The shape of the challenge is visible in Reflow's usage data:

Line chart showing Reflow token usage from early April through late June. Usage rises sharply during the workweek and falls on weekends, illustrating recurring peaks and valleys rather than steady utilization.
  • Reflow's token usage is very spiky: rising during weekdays, and falling on weekends. This is typical for enterprise products that see heavy usage during the work week, and minimal logins on nights and weekends.
  • For comparison, a flat GPU-hour capacity line would represent the fixed capacity Reflow needed to keep available under a rented-GPU model.
  • The difference between that fixed capacity and the actual usage shown here represents idle capacity Reflow avoided paying for with Parasail.

Reflow's uneven demand makes renting dedicated GPUs inefficient. The workload arrives in spikes, resulting in idle GPU hours when usage is low during off-hours.

The solution

Dedicated capacity, billed by the token

Reflow moved to Parasail's Dedicated Serverless — dedicated capacity billed per token — eliminating idle GPU costs by only paying for usage.

“Parasail helped us a lot with model optimization. We could test new models, benchmark them, and switch when we found improvements.”

Akbay TabakHead of Engineering, Reflow AI

Parasail also regularly helps the team benchmark new models, compare them against production needs, and adjust serving when a better fit emerges. In just the past seven days, Reflow tested nine distinct models across the Qwen and Llama model families.

The results

70% lower inference cost — with no hit to speed

Reflow lowered inference cost without giving up the responsiveness and reliability its product required.

“Overall, Parasail has been able to meet our aggressive performance requirements, while remaining stable enough that it runs in the background. I don't think about it much, which is what you want from infrastructure.”

Akbay TabakHead of Engineering, Reflow AI

What's next

Match the infrastructure to the shape of the workload

Reflow's story points to a broader lesson for AI-native startups: infrastructure decisions should follow workload shape. Some AI products are output-heavy, where tokens-per-second can define the user experience. Others are input-heavy, where cost, throughput, and end-to-end processing time matter more. Some workloads need real-time latency; some can move toward async or batch processing as the product matures.

Reflow's document ingestion and OCR workloads sit in the input-heavy category: they need to process a large amount of information efficiently, return useful results quickly, and avoid fixed infrastructure costs when demand is uneven and spiky. Dedicated GPUs can work when utilization is high and stable — but when the load is uneven, fixed capacity can force startups to pay for capacity that sits idle.

Start building today

Instantly run any open model — popular or specialized.