Build the runtime that schedules, serves and scales agent workloads for NiNI Cloud, measured against live traffic from the gateway.
NiNI Labs is the applied research lab in the family. Its first service, NiNI Cloud, is serverless inference for AI agents: a pipeline is imported as a graph of model and non-model stages, and the platform schedules, scales and traces it.
You will own pieces of that runtime: placement of requests across heterogeneous accelerators, scale-to-zero and resume without losing state, KV and prompt caches, and the tracing that attributes cost to every stage. The proving ground is real: the Nicrron gateway and the Ninoo agents generate the workloads, so ideas are judged on tail latency and dollars per useful token.
The button opens an email to careers@nicrron.ai with the subject line filled in. Tell us why this role and point us at one thing you have shipped. We reply to every application within a week.