Posted on: 21/05/2026
What You'll Do :
Data Engineering & Pipeline Development : Design and maintain backend data ingestion and embedding pipelines. Work hands-on with JupyterHub to set up environments, clone repositories, and run complex workflows.
Graph Database & Knowledge Management : Use Neo4jNeo4j to build graph-based solutions that map knowledge articles and workflows. Contribute to schema design, data integration, and performance tuning.
Data Quality & Optimization : Troubleshoot data quality issues, optimize Spark jobs, and manage partition sizes. Convert ECS direct connections to IO Meet using Spark and Python.
Authentication & Security : Solve JWT authentication challenges, research Kerberos for secure cluster connectivity, and manage credentials for Oracle DB and API clients.
DevOps & Collaboration : Work with GitLab for source control and Jira for project tracking. Participate in migrations from Azure DevOps to modern tools.
Your Tech Playground :
Programming & Tools : Python, JupyterHub, GitLab
Databases : Oracle DB, PostgreSQL + PGVector
Big Data : Spark, Iceberg, S3, Parquet, CSV
Graph & AI : Neo4j, LangChain, LLM integration
Security : HashiCorp Vault, Kerberos
Visualization : Power BI for dashboards and reporting
Skills That Make You Stand Out :
- Strong experience in data engineering data engineering, profiling, and ETL processes
- Expertise in Spark job tuning Spark job tuning and large-scale data exports
- Familiarity with embedding models embedding models like BGEM 3 and Nomic
- Ability to manage high-volume workloads high-volume workloads and optimize compute resources
- Knowledge of cloud and on-prem integration cloud and on-prem integration (Azure + on-prem systems)
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1638010