DATAREMOTECONTRACT

Data Engineer (Sports Analytics & AI)

Our client builds decision-making technology for American football, used by professional and college teams as well as by fans. Their models weigh player value, roster construction and in-game strategy, and every one of them is only as good as the data underneath it.

That data layer is where you come in. You will design and run the pipelines that carry everything from live game feeds to scouting notes and video metadata, and turn it into analytics-ready datasets that models, analysts and AI products can trust. You will work in the open with the machine learning and sports data teams rather than behind a ticket queue.

PythonSQLDatabricksSparkAirflowDelta LakeIcebergAWSKubernetesRAG

Responsibilities

  • Pipelines

    Build and operate ingestion, cleaning and transformation pipelines on Databricks, Airflow and Kubernetes.

  • Batch and streaming

    Write ETL and ELT workflows in Python and SQL for both scheduled and near real-time data.

  • Data for AI

    Shape datasets and tools so autonomous agents can query them safely, with evaluation and guardrails around AI-generated queries.

  • Retrieval

    Build RAG and vector search pipelines over structured statistics and unstructured material such as scouting notes and video metadata.

  • Modelling

    Design Delta, Parquet and Iceberg assets with versioning, lineage and reliability built in.

  • Orchestration

    Own scheduling, dependency tracking, monitoring and automated failure recovery.

  • Quality

    Enforce validation frameworks, schema contracts and audit logging so the numbers hold up.

  • Platform

    Evaluate tooling, set practice and tune performance and cost across compute, storage and orchestration.

Requirements

  • Experience: 3 to 8 years building and running production data pipelines as a data engineer or ETL developer.
  • Core stack: Strong Python and SQL, with hands-on Databricks, Spark or a comparable distributed processing framework.
  • Orchestration: Production use of Airflow, Dagster, Prefect or Luigi.
  • Foundations: Solid data modelling and warehousing, and comfort with modern lakehouse architectures.
  • Engineering practice: CI/CD, GitHub Actions, infrastructure as code and real testing for pipelines.
  • Cloud: AWS, GCP or Azure, plus Docker and Kubernetes.
  • AI tooling: Some exposure to LLM data work such as text-to-SQL, RAG, natural-language analytics or agent interfaces like MCP.
  • Collaboration: Able to work directly with machine learning, product and analytics people and explain trade-offs clearly.

What will be your next steps?

Quick non-technical conversation

Our initial conversation is a brief, non-technical discussion to understand your background and career aspirations. We're keen to learn about your communication style and how you approach teamwork and decision-making.

60 to 90 minutes technical interview

This in-depth technical assessment, lasting 60 to 90 minutes, is designed to evaluate your specific skills and expertise. We will present you with challenges relevant to our client’s requirements.

Client interview

In this stage, you will meet directly with the client for a final technical discussion. This interview will be similar in format to our internal technical assessment, allowing the client to see firsthand how your expertise aligns with their specific project needs and team.

Offer

Congratulations on successfully completing our evaluation process. We are pleased to extend an offer and recommend you to our clients.

Apply for this role

Fill in your details below. We'll get back to you shortly.

https://
Search and select skills...
Select applicable roles...

Enter your expected gross hourly rate in EUR.

PDF only, max 4 MB.