Back

LLM Platform Architecture Engineer — The Hong Kong Polytechnic University

Responsibilities

  1. Lead the architecture design, development, and implementation of LLM application platforms and data processing platforms; build data processing services for LLM training covering data collection, cleaning, governance, processing, and training dataset construction.
  2. Participate in the design, development, and deployment of MaaS (Model as a Service) platforms, productizing and serving different types of LLM capabilities to support internal business scenarios and external market expansion.
  3. Design and develop domain-specific LLM products, services, and solutions for different industries and business scenarios, covering model selection, technical architecture, Prompt Engineering, RAG, Agent, workflow orchestration, and system integration.
  4. Build Kubernetes-based LLM training, inference, and application service platforms, enabling containerized deployment, resource scheduling, elastic scaling, high availability, service governance, and operational monitoring.
  5. Participate in MLOps system development and improve the full lifecycle from data processing, experiment management, model training and evaluation, model registration, deployment and release, monitoring, to continuous iteration.
  6. Provide LLM technical consulting, solution design, and feasibility analysis based on customer and business requirements; identify business pain points and propose corresponding technical implementation recommendations.
  7. Design and develop LLM application prototypes, PoCs, and Demo systems to rapidly validate technical solutions and business value and move LLM solutions from proof of concept to real-world deployment.
  8. Collaborate closely with algorithm, product, software development, cloud platform, and business teams to develop, test, deploy, optimize, and continuously iterate LLM applications.
  9. Support algorithm teams in analyzing and assessing model fine-tuning requirements, including training data preparation, technical approach selection, fine-tuning tests, and effectiveness evaluation.
  10. Complete other related tasks assigned by supervisors.

Qualifications

  1. Master’s degree or above in Computer Science, Artificial Intelligence, Data Science, Business Analytics, or another related field.
  2. At least three years of relevant experience in AI, software development, cloud computing, IT services, product R&D, solution architecture, or pre-sales technical support.
  3. Familiarity with mainstream LLMs and multimodal models, including OpenAI GPT, Claude, Gemini, Qwen, and Wan, and understanding of core concepts such as Transformer, Token, Embedding, context windows, inference, and LLM training.
  4. Familiarity with LLM application development technologies including Prompt Engineering, Embedding, vector retrieval, RAG, Agent, Tool Calling, workflow orchestration, and model serving deployment.
  5. Familiarity with one or more mainstream AI application development frameworks, platforms, or tools, including LangChain, Dify, n8n, MCP, Skills, Claude Code, Alibaba Cloud Model Studio, or similar platforms.
  6. Familiarity with container and cloud-native technologies such as Docker and Kubernetes, with hands-on experience in one or more of the following areas:
  • Application deployment, operations, and troubleshooting in Kubernetes clusters;
  • Resource management for Deployment, StatefulSet, Service, Ingress, ConfigMap, Secret, and related objects;
  • Application deployment and configuration management tools such as Helm and Kustomize;
  • GPU resource scheduling, elastic scaling, high availability, and distributed training or inference environment construction;
  • Monitoring and observability tools such as Prometheus, Grafana, Loki, and OpenTelemetry.
  1. Familiarity with MLOps concepts, platform architecture, and engineering workflows; knowledge or hands-on use of one or more tools such as MLflow, Kubeflow, Airflow, Argo Workflows, KServe, Ray, or DVC.
  2. Experience with model lifecycle management, including data versioning, experiment tracking, model evaluation, model registration, deployment, canary release, version rollback, operational monitoring, and automated pipelines.
  3. Familiarity with CI/CD engineering practices and the ability to use tools such as GitLab CI, GitHub Actions, Jenkins, and Argo CD to build automated testing, image building, and deployment pipelines.
  4. Understanding of the basic concepts, common methods, and applicable scenarios of LLM fine-tuning, including full-parameter fine-tuning, LoRA, QLoRA, and PEFT; ability to support algorithm teams in assessing technical feasibility and conducting fine-tuning tests and effectiveness validation.
  5. Strong system architecture and engineering capabilities, with the ability to independently or jointly lead the full process of LLM application delivery from requirements analysis and solution design to prototype development and production deployment.
  6. Strong problem analysis, communication, coordination, and cross-functional collaboration skills, with the ability to translate business requirements into implementable technical solutions.

Preferred Qualifications

  1. Experience in LLM application deployment, AI product development, MaaS platform development, solution architecture, or AI pre-sales support.
  2. Experience building and operating production-grade Kubernetes clusters, large-scale GPU clusters, or AI computing platforms.
  3. Experience planning, building, or deploying MLOps platforms, with familiarity with automated pipelines for model training and inference workloads.
  4. Hands-on experience conducting model fine-tuning experiments using Hugging Face, Alibaba Cloud PAI, or similar platforms.
  5. Familiarity with LLM training data processing workflows and related technologies such as data cleaning, quality evaluation, deduplication, desensitization, annotation, and dataset management.
  6. Experience designing and developing multimodal models, Agent applications, RAG systems, or complex AI workflows.
  7. Familiarity with LLM inference frameworks and model serving technologies such as vLLM, TGI, Triton Inference Server, KServe, and Ray Serve.
  8. Good English communication skills and the ability to support cross-team, cross-region, or international business communication.
Deliver