Prilly
System Engineer - AI Engineer (100%)
- 29 September 2026
- 100%
- Permanent position
- Possible
About the job
We are looking for a System Engineer — AI Engineer to integrate open-source models and the surrounding ecosystem into production services running on-prem in our datacenters.
You'll join an experienced and enthusiastic team with a great team spirit, where knowledge is shared rather than hoarded. The services are business-critical, so the team shares a 24/7 on-call rotation.
System Engineer - AI Engineer (100%)
What you'll do
You own your solutions end to end — from understanding the business need to running them in production :
- Analyse the need with stakeholders, challenge it, and turn it into a design that fits our constraints — then implement, test, document, and deliver.
- Take ownership beyond the code: timelines, dependencies, communication with stakeholders, and the decisions that come with them.
- Evaluate, select, and deploy open-source models (LLMs, embeddings, rerankers, speech), and run the self-hosted inference stacks behind them (vLLM, Ollama) with sane quantization, batching, and model routing.
- Build the application layer — RAG pipelines, agents, tool calling, structured output — and own the data path behind it: ingestion, chunking, embeddings, vector stores, retrieval quality.
- Build evaluation harnesses and benchmarks so model choices are measured, not guessed.
- Integrate the ecosystem around the models: gateways, orchestration frameworks, model registries, observability and tracing.
- Tune for latency, throughput, and GPU utilization; handle capacity planning, scheduling, and production monitoring — and debug regressions when a model or version changes.
- Take part in the team's 24/7 on-call rotation (piquet), with compensation and time off in lieu — and improve the runbooks, alerting, and automation so quiet evenings are the norm, not the exception.
What you bring
- Master's degree in Computer Science or equivalent, with hands-on experience working with LLM-based systems.
- System engineering experience on Linux and Kubernetes.
- Strong Python; comfortable with production practices (testing, CI/CD, code review).
- Strong SQL skills.
- Experience self-hosting and serving open-source models.
- Working knowledge of retrieval, embeddings, and vector databases.
- Hands-on: not afraid of cabling a server or swapping a GPU.
- Autonomy and a sense of ownership — you follow a topic through to delivery rather than handing it over.
- Willingness to join a 24/7 on-call rotation, and the temperament to stay methodical under incident pressure.
- Pragmatism about open-source tooling — able to tell a solid project from a hyped one.
Nice to have
- GPU operations: drivers, CUDA, MIG, device plugins, scheduling on Kubernetes.
- Inference optimization (TensorRT-LLM, ONNX, speculative decoding, KV-cache tuning).
- Lightweight adaptation techniques (LoRA/QLoRA adapters on existing open-source models).
- Storage and networking design for AI workloads.
- Experience with AI safety, guardrails, and responsible deployment practices.
- Contributions to open-source AI projects.
What we offer
- Friendly and dynamic environment where proactivity and personal commitment are highly valued.
- Competitive salary and benefits system, very attractive pension fund conditions.
- Flexible working hours with 6 weeks of holidays.
- Modern and ergonomic workplace.
- Various fringe benefits.
- Free mobile subscription and discounts on our products.
- Possibility to widen your skills and experience due to a fast-moving and complex telco environment.
- A full open-source AI stack running on-prem in our datacenters — you own the hardware, not someone else's abstraction.
Always dreamed of having a job in a diverse and refreshing environment? Then we really need to meet. Please apply online (we only accept online applications with CV, employment references, diplomas).