
Data Engineer (RO)
Our Take
Engineer high-volume AV simulation data for a leading global automotive OEM at Spyrosoft.
What you’ll do
- Analyze sensor data for edge cases
- Create SQL, Python, Spark queries
- Process structured/semi-structured data
- Identify data for AV simulations
- Build data mining scripts & ETL
- Develop tools for analytics
What they’re looking for
- Advanced SQL, Python, Spark/PySpark
- Hands-on Databricks experience
- Over 4 years data engineering experience
- Experience with advanced data analytics
- Understanding of ML workflows
- Strong software engineering background
Skills & Focus Areas
- SQL
- Python
- Spark
- PySpark
- Databricks
- ETL
- Data mining
As posted by Spyrosoft
Requirements:
Core technical stack
SQL (advanced)
Python (advanced)
Spark / PySpark (advanced)
Hands-on experience with Databricks
Data engineering & analytics
Proven ability to write complex queries
Experience with advanced data analytics
Experience with time series analysis
ML & data knowledge
Understanding of ML workflows (data used for training, not an ML engineer role)
Software engineering
Strong software engineering background
Strong understanding of the implemented solutions
Experience & education
Over 4 years of commercial experience
University degree in Computer Science (nice to have)
Model of work:
B2B, Remote within Romania
Project description:
This long-term engagement supports simulation data processing for Autonomous Vehicle (AV) development. Key scenarios include obstacle detection, path planning, and complex traffic situations (e.g., tunnels, unusual vehicles, or temporary network issues).
The candidate will work with high-volume sensor data from a test AV fleet (8–12 cameras, LiDAR, radar), generating up to ~1TB/hour. Familiarity with the AV domain and real-world edge cases will be valuable. The team collaborates closely with a leading global automotive OEM to develop and validate safety-critical AV features.
The scope includes full-cycle data curation—from raw sensor input to simulation-ready datasets—and close cooperation with engineers and researchers.
Main responsibilities:
• Analyze real-world sensor data to identify edge cases (e.g., hard braking, nearby vehicles) • Create advanced SQL, Python, and Spark/PySpark queries for data filtering and transformation
• Work with internal tools for data search and auto-labeling workflows
• Process structured/semi-structured data (e.g., object detection output)
• Identify relevant data for AV simulations and ML pipelines
• Suggest and validate improvements in data discovery processes
• Build and maintain data mining scripts and ETL processes
• Develop tools to enhance analytics and streamline workflows