Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузите резюме
описание
Goldman Sachs is a global investment banking, securities and investment management firm that builds data foundations supporting AI and analytics capabilities.
задачи
Build, enhance and support batch and streaming data pipelines on the Lakehouse and AI data platform;
Refactor and modernise existing data flows to improve reliability, performance and maintainability;
Ensure data pipelines are production-ready, well tested and operationally supportable;
Develop raw, refined and curated datasets for analytics, reporting and AI use cases;
Apply data modelling principles to represent business entities, relationships and historical change accurately;
Work with consumers to shape usable, well-documented data products aligned with business needs;
Implement controls to validate data completeness, accuracy and consistency across pipelines and datasets;
Use reconciliation approaches to validate production outputs and investigate data breaks;
Contribute to standards for testing, monitoring and issue resolution;
Work with engineers, platform teams and data consumers to deliver agreed outcomes on time and to expected quality;
Communicate progress, risks, dependencies and design choices clearly;
Lead technical design across multiple datasets or pipeline domains;
Guide implementation standards, code quality and engineering practices within the team;
Lead delivery for a workstream, manage dependencies and support junior engineers.
требования
Bachelor’s or master’s degree in a relevant discipline, or equivalent practical experience;
Strong quantitative skills or data engineering expertise;
Strong hands-on programming experience in Python or Java;
Good working knowledge of SQL, including troubleshooting, optimisation and data analysis;
Ability to learn new tools, internal platforms and delivery workflows quickly;
Familiarity with version control, testing, release discipline and CI/CD practices;
Understanding of temporal data modelling, schema design, schema evolution and data compatibility;
Understanding of partitioning, clustering and other techniques for improving data performance at scale;
Ability to choose between normalised and denormalised models and between natural and surrogate keys;
Practical experience with data quality, reconciliation and root-cause analysis;
Experience building or supporting production data pipelines in a collaborative engineering environment;
Experience with distributed data processing frameworks such as Apache Spark;
Knowledge of JSON, Avro and Parquet;
Sound judgement in technical trade-offs and a structured approach to problem solving;
Willingness to work closely with stakeholders and partner teams;
Nice to have: Experience with ANSI SQL, Kafka, Snowflake, Apache Iceberg, Databricks, Hadoop ecosystem technologies, Sybase IQ, containerised or Kubernetes-based deployment approaches.
условия
Goldman Sachs provides training and development opportunities, firmwide networks, benefits, wellness and personal finance offerings, and mindfulness programs;
Reasonable accommodations are available during the recruiting process.