Senior Data Engineer
الوصف الوظيفي
About Intelmatix
Intelmatix is a cutting-edge deep tech company specializing in Artificial Intelligence (AI), established in July 2021 by a team of MIT scientists. Our mission is to revolutionize enterprises by transforming them into cognitive organizations—businesses that leverage AI and Decision Intelligence to drive smarter, more accurate, and error-resistant decision-making. By integrating advanced cognitive capabilities, we empower organizations to achieve superior outcomes across all facets of their operations, fostering innovation and sustainable growth.
The Role
We are seeking a seasoned Senior Data Engineer to spearhead the technical design, implementation, and delivery of an enterprise-grade Data Lakehouse, meticulously engineered to meet the demands of AI readiness. This pivotal role is instrumental in laying the groundwork for a transformative digital initiative, designed to support sophisticated AI agents, digital workforce solutions, and intricate knowledge graphs (Ontologies). The successful candidate will bring a wealth of experience in software development, with a strong emphasis on constructing and optimizing robust data pipelines, ensuring uncompromising data quality, and seamlessly integrating diverse data sources.
Based in Riyadh, you will be at the forefront of architecting a secure, scalable, and compliant data platform that adheres to stringent national data residency and cybersecurity regulations. Your responsibilities extend beyond system development; you will serve as a technical leader, mentoring client engineering teams through a collaborative "co-build" approach to cultivate long-term operational autonomy and expertise.
Key Responsibilities
- Lakehouse Architecture & Implementation: Design and deploy a unified Data Lakehouse leveraging the Medallion architecture (Bronze, Silver, Gold layers) and open table formats such as Delta Lake or Apache Iceberg, all hosted on cloud infrastructure within Saudi Arabia.
- Data Ingestion & Pipeline Engineering: Develop reusable, automated ingestion frameworks capable of processing both structured data (RDBMS, APIs) and unstructured data (PDFs, policy documents) to fuel downstream AI models and semantic reasoning engines.
- Data Quality & Governance: Implement automated data quality "circuit breakers" to validate completeness, uniqueness, and referential integrity, alongside comprehensive end-to-end data lineage tracking frameworks.
- Optimization: Enhance data processing workflows for peak performance, scalability, and cost-efficiency, ensuring optimal resource utilization and operational excellence.
- System Monitoring and Maintenance: Proactively monitor data systems, swiftly addressing Sev-1 incidents or other critical issues to maintain uninterrupted operations and reliability.
- Security & Compliance: Ensure the platform strictly complies with NCA (National Cybersecurity Authority) and NDMO (National Data Management Office) standards. Implement AES-256 encryption for data at rest, TLS 1.2+ for data in transit, robust Key Management Systems (KMS), and centralized audit logging.
- Access Control Integration: Design and deploy granular Role-Based Access Control (RBAC) and Attribute-Based Access Control (ABAC) systems, seamlessly integrating with existing enterprise Identity Providers such as Active Directory.
- Capability Building & Handover: Lead hands-on knowledge transfer sessions, engage in pair-programming with client engineers, develop operational runbooks, and conduct "Game Day" failure simulations to ensure the client’s team is fully equipped to independently operate the platform.
Qualifications & Experience
- Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field.
- Experience: Minimum of 5 years of hands-on experience in Data Engineering, Distributed Systems, or Big Data Architecture, with at least 2 years specifically leading Data Lakehouse or Cloud Data Platform implementations.
- Technical Skills & Core Technologies:
- Programming Languages: Proficiency in Python, Java, or Scala.
- Data Architecture & System Design: Strong expertise in designing data-intensive applications, complex data modeling, and schema design for enterprise environments.
- Distributed Systems & Lakehouse Technologies: Deep understanding of distributed computing, cloud-native architectures, and lakehouse technologies such as Delta Lake, Apache Iceberg, or Apache Hudi.
- Data Pipeline & ETL Tools: Experience with tools like Apache Spark, Apache Airflow, or similar frameworks for building scalable data pipelines.
- Cloud Platforms: Familiarity with cloud platforms such as AWS, Azure, or GCP, with a focus on deploying and managing data infrastructure.
- Security & Compliance: Knowledge of cybersecurity best practices, encryption methodologies, and compliance frameworks relevant to data governance.
يمكن أن يرتكب الذكاء الاصطناعي أخطاءً.