Skip to Content

Fullstack engineer (AI platform & APIs)

1 open position

Berget AI is building a developer-first, sovereign AI platform on open infrastructure. We run our own GPU hardware and provide Kubernetes-native inference APIs and platform services. We’re expanding the core product team and are looking for a backend developer who enjoys owning systems end-to-end and making complex infrastructure feel simple and reliable.

What you’ll work on

You will design, build, and operate backend services powering Berget AI’s platform, including inference APIs, OpenAI-compatible endpoints, internal product APIs, and usage and billing flows.

You’ll own features from idea to production: architecture, implementation, testing, deployment, monitoring, and iteration. You’ll keep releases flowing, fix bugs quickly, and continuously improve reliability and performance.

You’ll build Kubernetes-native services with strong focus on observability, security, and operational robustness. You’ll actively participate in incident response and ongoing security work, treating operational safety as a core product concern.

You’ll work closely with inference, platform, and community roles to translate real user needs into stable, developer-friendly APIs.

What you bring

You have strong experience building backend services in TypeScript/Node.js and designing clean, reliable APIs. We are using FluxCD as our GitOps platform and run everything i Kubernetes and we are obsessed with automating everything including testing, deployment so we can trust the system to heal itself when needed. You are used to using coding agents to accelerate your speed and accuracy.

You’re comfortable owning systems from PR to production and understand distributed systems fundamentals, failure modes, and secure service design.

You care about quality, reliability, and making infrastructure approachable for developers. Experience with Kubernetes-native application patterns is a strong plus. Familiarity with Python or inference runtimes (vLLM, SGLang) is a bonus, not a requirement.

You’re pragmatic, collaborative, and enjoy working in small teams with high ownership.

Why Berget

You’ll help shape the core APIs of a European sovereign AI platform, with real ownership, fast feedback loops, and close collaboration across engineering, product, and community.

Stockholm, Sweden
Full-Time

Inference & fine-tuning engineer (models, performance & security)

1 open position

Berget AI builds and operates sovereign AI infrastructure on our own GPU hardware. We run large-scale inference in production and are expanding into fine-tuning and post-training for real customer workloads. We’re looking for an inference engineer who masters model bring-up, performance tuning, and secure operation from model release to GPU.

What you’ll work on

You will rapidly evaluate, integrate, and bring new state-of-the-art models into production inference environments. This includes understanding model architectures, dependencies, runtime constraints, and quickly getting models running reliably on our GPU stack.

You’ll optimize inference runtimes and serving stacks (vLLM, Triton, SGLang, CUDA/ROCm), tuning metaparameters such as batching, parallelism, memory layouts, quantization strategies, and scheduling to maximize throughput, minimize latency, and control cost per token.

You will design and operate fine-tuning and post-training workflows: data preparation, training configuration, evaluation, model packaging, and safe rollout into production inference systems.

You’ll work extensively with caching strategies at multiple levels (model, KV/cache, request/result caching) to improve performance, efficiency, and isolation in multi-tenant environments.

You’ll extend Kubernetes-native orchestration for large-scale model serving, profile bottlenecks, benchmark improvements, and continuously harden reliability, observability, and security. Incident response, secure model handling, and controlled rollout are core parts of the role.

What you bring

You closely follow the latest open and commercial model releases and enjoy getting new models running fast in real systems.

You have hands-on experience with inference and ML runtimes such as vLLM, Triton, SGLang, CUDA or ROCm, and understand how model architecture, runtime configuration, and hardware interact.

You’re familiar with fine-tuning or post-training techniques and understand the operational and security implications of shipping trained models into production.

You’re comfortable working in Kubernetes-based environments, think practically about caching and isolation, and treat performance, reliability, and security as inseparable concerns.

Why Berget

You’ll have real ownership over how models are onboarded, optimized, and served in one of Europe’s most ambitious sovereign AI platforms. You’ll work close to hardware, platform, and product, with freedom to push performance and define best practices from day one.

Stockholm, Sweden
Flexible