HamburgerMenu
hirist

Job Description

What You'll Do :

Data Engineering & Pipeline Development : Design and maintain backend data ingestion and embedding pipelines. Work hands-on with JupyterHub to set up environments, clone repositories, and run complex workflows.

Graph Database & Knowledge Management : Use Neo4jNeo4j to build graph-based solutions that map knowledge articles and workflows. Contribute to schema design, data integration, and performance tuning.

Data Quality & Optimization : Troubleshoot data quality issues, optimize Spark jobs, and manage partition sizes. Convert ECS direct connections to IO Meet using Spark and Python.

Authentication & Security : Solve JWT authentication challenges, research Kerberos for secure cluster connectivity, and manage credentials for Oracle DB and API clients.

DevOps & Collaboration : Work with GitLab for source control and Jira for project tracking. Participate in migrations from Azure DevOps to modern tools.

Your Tech Playground :

Programming & Tools : Python, JupyterHub, GitLab

Databases : Oracle DB, PostgreSQL + PGVector

Big Data : Spark, Iceberg, S3, Parquet, CSV

Graph & AI : Neo4j, LangChain, LLM integration

Security : HashiCorp Vault, Kerberos

Visualization : Power BI for dashboards and reporting

Skills That Make You Stand Out :

- Strong experience in data engineering data engineering, profiling, and ETL processes

- Expertise in Spark job tuning Spark job tuning and large-scale data exports

- Familiarity with embedding models embedding models like BGEM 3 and Nomic

- Ability to manage high-volume workloads high-volume workloads and optimize compute resources

- Knowledge of cloud and on-prem integration cloud and on-prem integration (Azure + on-prem systems)

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...