Data Scientist , Engineering - Confluent

IBM
IBM

Data Science

Bengaluru, Karnataka, India

Posted on Sep 16, 2026
Introduction

At IBM Software, we transform client challenges into solutions. Building the world’s leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You’ll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM’s product and technology landscape. Here, you’ll have the tools and opportunities to advance your career while creating software that changes the world.

Your role and responsibilities

About the Role

As a Data Engineer II of Confluent, an IBM company, Data team you will own meaningful slices of our data platform end-to-end -- from ingestion through transformation to the data products that power decision-making across the company. You will be part of Confluent R&D team and work closely with various business stakeholder like Prod&Eng, Marketing, FieldsOps, Sales etc to address their analytical needs. You would be working closely with the overall IBM CDO and CIO team to enable the wider organisation with data needs from the Confluent ecosystem. You'll design and operate batch and real-time pipelines, partner closely with Data Scientists, Analysts, and business stakeholders, and raise the bar on craftsmanship, reliability, and developer productivity for the team.

You'll work in a modern, AI-augmented engineering environment: shipping faster with Claude Code and other AI coding assistants, building AI-powered internal tools (natural-language-to-SQL, automated data quality, lineage, anomaly detection), and bringing GenAI thoughtfully into data workflows. We expect you to take significant projects from ambiguous problem statement to delivered outcome, mentor newer engineers, and represent the team confidently with cross-functional partners.

What You Will Do

* Design, build, and operate efficient, reliable, and well-documented data pipelines across batch and streaming systems -- from source ingestion through the warehouse to consumer-facing data products.
* Own the data models, SLAs, and quality contracts for the domains you cover; treat documentation, testing, and observability as first-class deliverables.
* Improve the data platform itself: introduce or extend tooling, templates, CI/CD, testing frameworks, and monitoring that make the entire team faster and more reliable.
* Adopt and champion AI-assisted engineering -- use Claude Code and similar tools as part of your daily workflow, build internal AI-powered utilities (text-to-SQL on the semantic layer, automated PR review, log triage, data discovery), and evaluate where LLMs add durable value vs. hype.

* Partner directly with Data Scientists, Analysts, and business stakeholders to translate ambiguous requirements into well-scoped deliverables; communicate trade-offs, anticipate push-back, and drive alignment without needing escalation.
* Mentor interns, new hires, and L2 engineers -- through onboarding, code review, design feedback, and pairing.
* Contribute to cloud infrastructure and governance: IAM, service accounts, secrets management, cost optimization, and Terraform-managed GCP / Confluent Cloud resources.
* Drive incident response and root-cause analysis for the pipelines you own; close the loop with durable fixes, runbooks, and prevention.

Required education
Bachelor's Degree
Preferred education
Master's Degree
Required technical and professional expertise

6 to 9 years of experience in Data Engineering at a technology company, including production ownership of non-trivial pipelines and data models.
* Strong command of advanced SQL and Python -- you can write production-grade code, review others' code, and pick the right tool for the job.
*Have end to end ownership of Data pipeline and a clear grasp on requirement gathering, data modeling, data quality and access management
*Solid grounding in dimensional modeling, data warehousing, and ELT/ETL fundamentals; you can design a clean star schema and explain the trade-offs.
* Hands-on experience operating orchestration frameworks in production (Airflow / Cloud Composer, Dagster, or similar).
* Hands-on experience building streaming and event-driven pipelines -- Kafka, Flink, or equivalent -- and reasoning about exactly-once, schema evolution, and back-pressure.
* Experience on a major cloud platform (GCP preferred) and with infrastructure-as-code (Terraform).
* Demonstrated ability to operate with moderate ambiguity: scope your own work, identify what to clarify, recommend an approach, get buy-in, and ship.
* Fluency with AI coding assistants (Claude Code, Cursor, Copilot) as a daily tool -- and judgment about where they help vs. where they hurt.
* Clear written and verbal communication; comfortable presenting designs, post-mortems, and trade-offs to technical and non-technical audiences.

Preferred technical and professional experience

Experience with open table formats (Apache Iceberg, Delta Lake) and modern lakehouse patterns.

- Familiarity with data quality / observability tooling (Great Expectations, Soda, Elementary, Monte Carlo) and catalog / lineage systems (DataHub, Atlan, Unity Catalog).

- Experience building or integrating with AI/ML infrastructure: feature stores (Feast, Tecton), vector databases (pgvector, Pinecone, Weaviate), or RAG pipelines for internal knowledge systems.

- Experience contributing to or operating a semantic layer that powers self-serve analytics and natural-language interfaces.

- Familiarity with Confluent Cloud, Kafka Connect, and Flink SQL.

- Open-source contributions or technical writing in the data engineering space.

Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience.

Years of Experience:
6 - 9

ABOUT BUSINESS UNIT

IBM Software infuses core business operations with intelligence—from machine learning to generative AI—to help make organizations more responsive, productive, and resilient. IBM Software helps clients put AI into action now to create real value with trust, speed, and confidence across digital labor, IT automation, application modernization, security, and sustainability. Critical to this is the ability to make use of all data, because AI is only as good as the data that fuels it. In most organizations data is spread across multiple clouds, on premises, in private datacenters, and at the edge. IBM’s AI and data platform scales and accelerates the impact of AI with trusted data, and provides leading capabilities to train, tune and deploy AI across business. IBM’s hybrid cloud platform is one of the most comprehensive and consistent approach to development, security, and operations across hybrid environments—a flexible foundation for leveraging data, wherever it resides, to extend AI deep into a business.

YOUR LIFE @ IBM

In a world where technology never stands still, we understand that, dedication to our clients success, innovation that matters, and trust and personal responsibility in all our relationships, lives in what we do as IBMers as we strive to be the catalyst that makes the world work better.

Being an IBMer means you’ll be able to learn and develop yourself and your career, you’ll be encouraged to be courageous and experiment everyday, all whilst having continuous trust and support in an environment where everyone can thrive whatever their personal or professional background.

Our IBMers are growth minded, always staying curious, open to feedback and learning new information and skills to constantly transform themselves and our company. They are trusted to provide on-going feedback to help other IBMers grow, as well as collaborate with colleagues keeping in mind a team focused approach to include different perspectives to drive exceptional outcomes for our customers. The courage our IBMers have to make critical decisions everyday is essential to IBM becoming the catalyst for progress, always embracing challenges with resources they have to hand, a can-do attitude and always striving for an outcome focused approach within everything that they do.

Are you ready to be an IBMer?

ABOUT IBM

IBM’s greatest invention is the IBMer. We believe that through the application of intelligence, reason and science, we can improve business, society and the human condition, bringing the power of an open hybrid cloud and AI strategy to life for our clients and partners around the world.

Restlessly reinventing since 1911, we are not only one of the largest corporate organizations in the world, we’re also one of the biggest technology and consulting employers, with many of the Fortune 500 companies relying on the IBM Cloud to run their business.

At IBM, we pride ourselves on being an early adopter of artificial intelligence, quantum computing and blockchain. Now it’s time for you to join us on our journey to being a responsible technology innovator and a force for good in the world.

IBM is proud to be an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, neurodivergence, age, or other characteristics protected by the applicable law. IBM is also committed to compliance with all fair employment practices regarding citizenship and immigration status.

OTHER RELEVANT JOB DETAILS

When applying to jobs of your interest, we recommend that you do so for those that match your experience and expertise. Our recruiters advise that you apply to not more than 3 roles in a year for the best candidate experience. For additional information about location requirements, please discuss with the recruiter following submission of your application.