
We're not tokenmaxxers, we want the most intelligence per dollar, and Inference Router routes every request to the model that gets us the best result.
Travis Beale
CTO, Consumable
Consumable is a fast-growing adtech company that creates unique advertising experiences across audio, video, mobile, and web. With a real-time platform handling up to one million queries per second at peak times, performance, scalability, and simplicity are mission-critical.
CTO Travis Beale and his team consolidated Consumable’s compute, load balancing, and databases onto DigitalOcean in 2021. Since then, he’s worked to integrate additional AI offerings, including Signal Boost (an inline LLM enrichment to fill in missing program metadata) and nano models tuned for low-reasoning JSON metadata. For Beale, the payoff has been predictable scaling, costs he can forecast, and the freedom to change models without rebuilding.
When Beale joined Consumable, the engineering team needed to integrate tech stacks and hosting environments from various acquisitions. A recently acquired product, ServerBid, was already running on DigitalOcean while other systems were hosted on AWS.
“We had to make a choice between consolidating our server resources in DigitalOcean or consolidating them in AWS. And based on our cost modeling, DigitalOcean was the better deal for our workload,” Beale says.
That migration decision set the foundation for a broader shift to DigitalOcean’s cloud infrastructure, where they could better manage costs while meeting the company’s performance demands and growth philosophy. Beale explains that Consumable is a bootstrapped company that doesn’t want to speculatively build out infrastructure and then be trapped once customer demand changes.
“We want to build infrastructure that follows growth as it arrives and is as flexible as we are. I want to avoid rigid systems that trap us when customer demand changes. We’re not tokenmaxxers, and Inference Router lets us route every request to the model that gets us the best result," Beale says.
To consistently meet infrastructure demand and avoid overforecasting, Beale works with Senior Technical Account Manager Raph Sirvent to discuss what might be needed over the next three months, based on how the business is performing.
From the outset, pricing clarity played a key role in Consumable’s decision. Beale noticed immediate cost savings when compared to running the same workload on AWS.* Even as their infrastructure needs evolved and they adopted Premium Droplet® instances to address latency issues, DigitalOcean’s team worked closely with them to maintain a manageable cost structure and to design infrastructure to support the necessary performance.
They faced a similar situation when running AI models on DigitalOcean. Beale’s team had worked with OpenAI before, but was running into billing predictability issues, which kept appearing as miscellaneous $40 charges. Now, the costs associated with models just appear as another line item in their monthly cloud invoice.
"The consolidation on our existing DigitalOcean bill made things easier for me to track and easier for our finance team. The predictability aspect was really key,” Beale says.
In programmatic advertising, milliseconds matter. Consumable runs a high-throughput, low-latency platform where response times must remain under a few hundred milliseconds.
Since DigitalOcean is designed to support consistent, scalable infrastructure, the Consumable team began exploring it to run their AI models instead of going through OpenAI directly. By Beale’s account, his team’s spend level put them under daily rate limits, making it much harder to experiment with and test features.
“We would get into a catch-22 situation—we had demand to scale the product more, but because we weren’t a big enough customer, our rate limit was capped, and we couldn’t use more. You’re rate-limited because you haven’t grown enough. But we can’t grow because we’re being rate-limited. We repointed our products at DigitalOcean, and that problem just went away,” Beale said.
With higher capacity available on DigitalOcean Serverless Inference, the team experiences less friction when forecasting capacity and fewer trade-offs in project priorities due to limited resources.
“Travis has the headroom he needs on inference; performance has held up, and he isn’t locked into one model or one vendor. That fits how he thinks about infrastructure,” says Sirvent, Senior Technical Account Manager.
With Consumable using OpenAI, one main concern was keeping up with how quickly models change. Beale started looking into other options when the vendor announced it would sunset one of its models, requiring developers to migrate to a newer model on a deadline, leaving them back at square one for testing and validation.
“Model sunsets are the real pain point: when a vendor retires a model, teams can be left re-validating everything on a deadline. Going forward, if they get word of a model sunset, they can compare performance via Evaluations and then seamlessly add the new model into the routing pool with Inference Router, instead of guessing," says Sirvent, Senior Technical Account Manager.
Beale explains this is the new cost of AI-enabled apps; development teams can spend so much time on one model with regression testing and additional engineering resources, only to have 60 days to move off of it and start the testing process all over again.
On DigitalOcean, not only can Consumable select from a variety of models as part of the model catalog, but developers can use Inference Router to effectively select the right model for each task. This can reduce reliance on a single model and can lower production costs.
“We run a real-time platform that peaks at a million queries per second, so the last thing we want is to be dependent on a model that is deprecated within 6 months of launch. DigitalOcean Inference Router lets us put multiple models in the mix and route each request to the one that gets us the best result, without redoing our testing every time a model changes. We don’t have anything to prove about how much AI we use. We make intelligent and optimal use of it, and that’s how we get the most intelligence per dollar,” Beale says.
Consumable uses multiple DigitalOcean products for its platform, including Droplet instances, Load Balancers, Managed Databases, App Platform, Kubernetes (DOKS), Container Registry, command-line tools for automation and monitoring, and, most recently, Serverless Inference and Inference Router for AI workloads.
Beale acknowledged that although migrating away from AWS was a major undertaking, it was relatively seamless.
“There were some things that we were using on AWS that were a managed offering—our Redis cluster was one—so we just self-hosted that,” he says.
He noted that while building and tuning their own Redis instance was more work, it gave them full control over performance. Migrating infrastructure may not be something anyone wants to do every year, but it was a move that resulted in many net positives for the company.
That same ease reappeared when Consumable moved its AI workloads to the DigitalOcean Inference Router.
"The move was very easy. It’s the same API—everybody’s standardizing around the OpenAI API. We picked a candidate product from our side, made a quick two-minute code change, and it worked. We did see some scaling challenges with increased latency, but it was solved in 24 hours,” Beale says.
As Consumable prepares to expand its services, Beale is also testing AI as a part of this initiative. He knows it’s not a passing trend, but a must-have to remain competitive and meet customer demand.
Thinking about what that growth will require, Beale predicts a smooth scale-up with DigitalOcean. Because his team can easily spin up new projects without a capacity conversation, they can freely test features and model performance. As Beale puts it, that’s "going to make DigitalOcean our vendor of choice for the larger products.”
“I don’t see us creating any new products that don’t incorporate a large language model in some way, and there really aren’t any of our legacy products that aren’t good candidates for incorporating an LLM. We haven’t done that with all of them yet, but it’s only a matter of time," he explains.
* Results in customer environments may vary depending on configuration, implementation, and usage. Results and/or savings are not guaranteed.
DISCLAIMER: Any references to third-party companies, trademarks, or logos in this document are for informational purposes only and do not imply any affiliation with, sponsorship by, or endorsement of those third parties.

Read how DigitalOcean and NVIDIA help Hippocratic AI create safe, compliant AI agents for healthcare appointment management and close care gaps.

Traversal uses DigitalOcean’s AI and GPU infrastructure to power advanced root cause analysis, helping enterprises understand complex system issues.

From launch spikes to daily play, Double Eleven trusts DigitalOcean Droplets to keep Rust online.
From GPU-powered inference and Kubernetes to managed databases and storage, get everything you need to build, scale, and deploy intelligent applications.
