Meridian Robotics
tier A- Serves models on self-hosted vLLMvLLMAn open-source engine for serving large language models fast. Self-hosting it means running models on your own GPUs.confirmedconfirmed: engineering blog · may 12
Serving cost and latency are their team's problem directly. That is the budget line you are selling into.
- Two open roles for inferenceinferenceRunning a trained AI model to get answers. This, not training, is where most of the GPU bill goes. optimizationconfirmedconfirmed: careers page · may 18
They are trying to hire their way out of this. A pilot that lands before those hires ramp is cheaper and faster, and they know it.
- GPUGPUThe chips AI models run on. Renting them is one of the biggest line items on an AI company's bill. spend flagged in Q2 planningconfirmedconfirmed: discovery call transcript · may 20
The pain has a named line in a planning doc and exec attention this quarter. Deals move when that is true.
- Evaluating managed endpointsmanaged endpointsSomeone else hosts the AI model and you call an API. The opposite of running it yourself.inferredinferred: founder forum thread · may 15
They are already shopping. Assume at least one competitor call has happened; speed matters more than polish now.
This deal is about serving cost at scale, and the person who feels it is the infra lead, not procurement. They have exec attention on the GPUGPUThe chips AI models run on. Renting them is one of the biggest line items on an AI company's bill. bill and hiring underway, which means they will fix this with or without you inside two quarters. The way in is a benchmark on their own traffic that beats what their new hires could build by summer.
model judgment · labelled, never mixed with the evidence above- Priya NairHead of Infrastructurestart here
Owns the GPUGPUThe chips AI models run on. Renting them is one of the biggest line items on an AI company's bill. bill and wrote the blog post you will be quoting. Technical, direct, allergic to marketing language. Open with the idle-replica number from her own call.
- Dan OkaforVP Engineeringnext
Final sign-off. He does not care what it costs; he cares that p95p95The response time the 95th-slowest request gets. The number engineers watch to know users are not waiting. never regresses. Bring the latency chart, not the pricing page.
- Procurementnot yet engagedhold
Correctly out of the loop for now. Bringing them in before the pilot is scoped adds four weeks for nothing.
“We're burning half our GPUGPUThe chips AI models run on. Renting them is one of the biggest line items on an AI company's bill. budget on idle replicasidle replicasModel servers kept running with no traffic, just in case. Paid by the hour, used never..”
“If p95p95The response time the 95th-slowest request gets. The number engineers watch to know users are not waiting. goes over 400ms, product will notice before we do.”
- they say
“We can just build this ourselves.”
you sayAgree that they could. Then price the engineering months against the pilot cost and offer the benchmark on their own traffic. The comparison does the arguing.
- they say
“Managed endpointsManaged endpointsSomeone else hosts the AI model and you call an API. The opposite of running it yourself. lock us in.”
you sayDeploy in their VPCVPCA private, walled-off slice of a public cloud. Where careful companies keep their data.. Weights and routing configs stay theirs, and the export path is shown up front, not promised.
- they say
“Not this quarter.”
you sayThe GPUGPUThe chips AI models run on. Renting them is one of the biggest line items on an AI company's bill. bill is already flagged for Q2 planning. A two-week pilot lands before planning closes; waiting is the expensive option.
Thanks for the numbers. Before the demo, can you clarify how autoscalingautoscalingCapacity that grows and shrinks with traffic automatically, instead of being provisioned by hand. handles our burst traffic? Evenings spike hard for us.
Short version: scale-to-zeroscale-to-zeroServers shut off completely when idle, so nothing is billed for waiting. on idle replicasidle replicasModel servers kept running with no traffic, just in case. Paid by the hour, used never., warm poolwarm poolA few servers kept ready so a traffic spike does not hit a cold start. for burst. Your May post mentions 3x evening peaks; we would size the warm poolwarm poolA few servers kept ready so a traffic spike does not hit a cold start. off that curve and show it live in the demo.
drafted by gtminfra from the sources above · review before send- Run the demo against a sample of their traffic, not a toy model
- Send Priya the pilot scope with the warm-pool sizing
- Ask for the latency SLOSLOService level objective: the response-time promise an engineering team commits to keeping. on the demo call; it is still unknown
- blogServing LLMs at Meridianmay 12
- calldiscovery call · 34 min · recordedmay 20
- careersinference optimization rolesmay 18
- forumfounder thread on managed endpointsmay 15
- crmdeal record · stage 2 notesmay 21
- hubspotdeal record2m ago
- gongcall recording5m ago
- gmailthread: benchmark numbers2m ago
- notionaccount notes10m ago
Every account in your ICP gets a page like this, at yourcompany.gtminfra.com.
Book a demo →