About the role
At Ocean Infinity, we're on a bold mission to subsea data using innovative technology.
Using cutting-edge robotics, autonomous technology and world-class software, we're transforming how complex operations are carried out at sea. Our uncrewed systems are making ocean exploration and operations safer, smarter and more sustainable, helping unlock the vast potential of our oceans while reducing environmental impact. 🌍
This isn't a vision for the future. It's happening right now.
As we continue to push the boundaries of innovation and redefine what's possible in the maritime industry, we're looking for exceptional talent to join our fast-growing team.
Ocean Infinity is seeking a motivated and technically capable Data Engineer to join our team. The successful candidate will focus on building and implementing scalable data pipelines and automation solutions that transform complex manual workflows into reliable, production-grade data systems.
You will contribute to delivering robust, production-ready data infrastructure, working across data, product, and engineering teams to apply engineering standards and provide high-quality curated datasets that power analytics, operational systems, and AI initiatives across the company.
What you will do! 🚀
- Develop and implement scalable data pipelines that automate manual data processes and improve reliability and efficiency across the organization.
- Build and maintain end-to-end data transformation workflows (Bronze → Silver → Gold), focusing on implementing business logic, data validation, and reusable transformation components.
- Develop robust Python-based data processing solutions, including scripts, services, and automation tools that support ingestion, transformation, and data delivery.
- Implement data integration pipelines across a variety of sources, including operational databases, external APIs, and event/streaming data, using both batch and near-real-time approaches.
- Develop and maintain data quality checks, validation rules, and testing frameworks to ensure accuracy, consistency, and reliability of data products.
- Build backend data services and lightweight APIs that expose curated datasets to analytics tools, reporting systems, and downstream applications.
- Collaborate closely with data analysts, product teams, and AI engineers to translate data requirements into efficient, production-ready implementations.
- Optimize and refactor existing pipelines and processes to improve performance, reliability, and maintainability, with a strong focus on code quality and automation.
- Follow and contribute to engineering best practices, including version control standards, modular pipeline design, CI/CD for data workflows, and documentation of implemented solutions.
- Participate in code reviews and pairing, learning from more experienced engineers and sharing implementation knowledge and reusable patterns with the team.
What we look for!🔍
- A degree in Computer Science, Mathematics, Engineering, or a related field, or equivalent practical experience.
- Minimum 2 years of experience as a Data Engineer, Backend Engineer, or Software Engineer working with data-intensive systems and data pipelines.
- Solid software engineering skills in Python, with experience writing reliable, maintainable data processing code and automation workflows.
- Good SQL skills, with an understanding of query optimization and working with analytical datasets.
- Experience building and maintaining data pipelines and transformation workflows, ideally in cloud-based environments.
- Hands-on experience with at least one cloud platform (e.g., AWS, GCP, or Azure), particularly for building and running data pipelines, storage systems, and compute workloads.
- Familiarity with data lake and lakehouse environments, including object storage systems (e.g., S3 or equivalent) and structured transformation layers (e.g., Bronze/Silver/Gold patterns).
- Understanding of unstructured and semi-structured data storage patterns (e.g., JSON, logs, event data, files in object storage) and how to process them efficiently.
- Experience with relational and/or NoSQL databases (e.g., Postgres, MongoDB, Redis), including schema design and indexing.
- Working knowledge of containerization technologies such as Docker.
- Good understanding of data modelling principles and how to structure datasets for analytics and downstream consumption.
- Exposure to distributed systems concepts and large-scale data processing patterns, with emphasis on practical implementation and debugging.
- Strong problem-solving skills with a focus on reliability, performance, and automation in production systems.
- Ability to collaborate effectively with analysts, product teams, and engineers, translating requirements into robust, testable implementations.
- Interest in improving and refactoring existing pipelines and workflows to increase reliability, maintainability, and automation.
- Comfortable working in modern engineering practices including Git-based workflows, code reviews, CI/CD pipelines, and automated testing for data systems.
- Exposure to workflow orchestration tools such as Airflow, Prefect, or similar, including scheduling, dependency management, retries, and observability.
Nice to have! ✨
- Experience working with telemetry, sensor, location, and operational technology (OT) data workloads.
- Experience processing and managing large-scale video, image, and unstructured media datasets, including streaming ingestion and analytics workflows.
- Understanding of feature engineering patterns and data preparation workflows supporting machine learning systems.
- Experience working with maritime, fleet, vessel operations, logistics, or industrial operational domains, including telematics, tracking, asset monitoring, or operational analytics use cases.
Benefits
*Salary: London £65,000 / Porto up to 55,000 EUR Per Annum*
Source: the employer's own careers page.