Back

On-Device Training & Inference Infrastructure Engineer

Responsibilities

  1. Participate in the training, fine-tuning, and experimental iteration of on-device/local LLMs and multimodal models, and drive deployment on real devices.
  2. Optimize on-device model inference, including quantization, distillation, KV Cache, long-context optimization, and inference pipeline tuning.
  3. Participate in the end-to-end workflow of local training, inference, and deployment, connecting model training through device-side execution.
  4. Conduct performance analysis and optimization on CPU, GPU, and NPU platforms, focusing on latency, throughput, memory usage, and power consumption.
  5. Participate in the development of on-device runtimes, inference frameworks, and toolchains to improve model stability and performance.
  6. Collaborate with algorithm, systems, and product teams to bring on-device AI capabilities into production.

Qualifications

  1. Bachelor’s degree or above in Computer Science, Artificial Intelligence, Software Engineering, or a related field.
  2. Strong engineering skills and proficiency in Python, with good coding practices and debugging capabilities.
  3. Proficiency in PyTorch, with hands-on model training or fine-tuning experience beyond API-only usage.
  4. Familiarity with Linux development environments and the ability to independently diagnose issues, analyze logs, and reproduce experiments.
  5. Hands-on project experience in at least one of the following areas:
  • LLM training/fine-tuning (LoRA / QLoRA / multimodal)
  • Inference optimization (quantization / KV Cache / latency optimization)
  • Model deployment / on-device deployment

Preferred Qualifications

  1. Experience with on-device/local AI projects on PCs, mobile devices, or edge devices.
  2. Familiarity with local training, inference, or deployment toolchains such as Unsloth, KTransformers, Nexa AI / Nexa SDK.
  3. Hands-on experience with LoRA, QLoRA, PTQ, QAT, distillation, long-context optimization, or KV Cache.
  4. Familiarity with CUDA, Triton, C++, or other performance optimization tools, with kernel or inference acceleration experience.
  5. Experience with local model export, cross-platform deployment, or runtime adaptation, and understanding of the full process from training to device-side delivery.
  6. Open-source contributions or demonstrable GitHub projects, technical blogs, or similar work.
Deliver