AI Engineer
الوصف الوظيفي
AI Engineer – Drive Innovation in Large-Scale Language Model Solutions
Are you passionate about shaping the future of artificial intelligence through cutting-edge, production-grade systems? We are seeking a highly skilled AI Engineer to lead the design, development, and deployment of advanced AI solutions, including fine-tuned large language models (LLMs) and multi-agent systems. This role demands a blend of technical expertise, strategic vision, and leadership to architect scalable, secure, and high-performance AI platforms that deliver measurable business impact.
In this dynamic position, you will collaborate with cross-functional teams to translate complex business requirements into robust, deployable AI prototypes. Your work will span the entire AI lifecycle—from model fine-tuning and evaluation to system integration, observability, and governance—ensuring seamless scalability and compliance with enterprise standards. You will also play a pivotal role in mentoring and guiding a team of cloud and AI operations engineers, fostering a culture of excellence in software development, DevOps, and MLOps practices.
Key Responsibilities
Your responsibilities will include, but are not limited to:
- AI System Development and Optimization: Build, refine, and evaluate domain-specific LLM systems using advanced techniques such as QLoRA and PEFT on open-weight models like Llama-3 and Mistral. Design reproducible evaluation frameworks and A/B testing methodologies to measure performance metrics, including task success rates, safety compliance, and latency distributions (p50/p95).
- Multi-Agent and RAG System Architecture: Develop sophisticated multi-agent systems and Retrieval-Augmented Generation (RAG) pipelines using tools like LangGraph, FastAPI, and vector databases. Scale these systems from prototype to production, ensuring high availability, reliability, and performance.
- Safety and Compliance Guardrails: Implement rigorous input/output validation, allowlist/denylist policies, and automated controls to mitigate risks associated with model outputs. Ensure alignment with organizational safety standards and regulatory requirements.
- Prototyping and Stakeholder Engagement: Translate business use cases into functional prototypes with clear acceptance criteria. Present and demonstrate these solutions to stakeholders, driving adoption and continuous improvement.
- Cloud and MLOps Infrastructure: Design and manage cloud-based AI infrastructure on platforms like Azure, OCI, or GCP, leveraging Kubernetes and containerized runtimes. Optimize MLOps workflows, including CI/CD pipelines, GitOps-based deployments (Argo CD), and automated release promotion across environments.
- Observability and Security: Establish end-to-end observability using tools like Azure Monitor, Application Insights, and ELK. Implement network security baselines, automated code quality scanning (SonarQube, Black Duck), and compliance checks to ensure robust system resilience.
- Disaster Recovery and High Availability: Lead the design of disaster recovery strategies, including automated backups, failover mechanisms, and documented recovery time objectives (RTO) and recovery point objectives (RPO).
- Engineering Leadership: Mentor and lead a team of cloud/AI operations engineers, defining best practices for monitoring, incident response, and release governance. Standardize SDLC processes, including branching strategies, PR governance, and release management, to enhance delivery efficiency.
- Tooling and Workflow Optimization: Consolidate engineering tooling and workflows, driving migrations and platform standardization to eliminate fragmentation and accelerate delivery. Develop comprehensive handover documentation and runbooks for operational continuity.
- Vendor and Licensing Negotiations: Support negotiations for cloud enterprise agreements and licensing, ensuring cost-effectiveness and alignment with organizational goals.
Qualifications and Experience
To excel in this role, you must possess:
- Education: A Bachelor’s degree in Software Engineering, Computer Science, or a related field. A Master’s degree in Applied AI, Machine Learning, or a related discipline is highly preferred.
- Technical Proficiency: 6–8+ years of experience in software, DevOps, or platform engineering, with at least 2 years in applied AI or ML engineering. Proven track record of delivering production AI/LLM systems—beyond research or notebook-stage experimentation.
- Programming Languages: Proficiency in Python with experience in Bash and YAML. Familiarity with PyTorch and the Hugging Face Transformers library is a significant advantage.
- Cloud and Containerization: Hands-on experience with Kubernetes, Docker/Podman, and Terraform. Strong production experience with at least one major cloud provider, with a preference for Azure (OCI or GCP also acceptable).
- CI/CD and GitOps: Demonstrated ownership of CI/CD pipelines at scale, including Azure DevOps, GitHub Actions, and GitOps release models like Argo CD.
- Leadership and Collaboration: Experience leading a team and establishing engineering standards across multiple engineering squads. Ability to drive cultural change and improve lead time and deployment frequency.
- Preferred Specializations: Fine-tuning experience with QLoRA/LoRA on GPU clusters. Expertise in vector databases (Milvus, Pinecone, or Weaviate) and RAG retrieval design. Familiarity with delivering solutions for large-scale national digital platforms, including compliance with local standards.
- Language Proficiency: Professional proficiency in both Arabic and English.
Technical Environment
You will work within a robust technical stack, including:
- Python, FastAPI, PyTorch, Transformers, LangGraph
- Vector databases (Milvus, Pinecone, Weaviate)
- Redis, PostgreSQL
- Kubernetes, Docker/Podman, Terraform
- Argo CD, Azure DevOps, GitHub Actions
- Azure ML, Azure Monitor, Application Insights, ELK
- SonarQube, Black Duck, Fortinet FW/WAF
This is an opportunity to make a lasting impact on AI-driven innovation while leading a team of talented engineers toward excellence. If you are ready to take on this challenge and drive transformative AI solutions, we invite you to apply.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.
ℹ️ إخلاء مسؤولية توظيف:
موقع وظائف السعودية (ksajobshub.com) هو محرك بحث ومجمع لإعلانات الوظائف من المصادر والشركات الرسمية في المملكة العربية السعودية. نحن لا نتقاضى أي مبالغ مالية أو رسوم من الباحثين عن عمل، وتتم عمليات التقديم مباشرة عبر الانتقال للرابط الأصلي للجهة المعلنة.