Upload Button Icon Add office photos
filter salaries All Filters

171 Nvidia Jobs

Senior Site Reliability Engineer, Data Science and ML Platforms

5-8 years

Hyderabad / Secunderabad, Pune, Gurgaon / Gurugram + 1 more

1 vacancy

Senior Site Reliability Engineer, Data Science and ML Platforms

Nvidia

posted 10hr ago

Job Description

Are you passionate about building and maintaining large-scale production systems that support advanced data science and machine learning applications? Do you want to join a team at the heart of NVIDIA's data-driven decision-making culture? If so, we have a great opportunity for you! NVIDIA is seeking a Senior Site Reliability Engineer (SRE) for the Data Science & ML Platform(s) team. The role involves designing, building, and maintaining services that enable real-time data analytics, streaming, data lakes, observability and ML/AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of the platform, as well as applying SRE principles to improve production systems and optimize service SLOs. Additionally, collaboration with our customers to plan implement changes to the existing system, while monitoring capacity, latency, and performance is part of the role.

To succeed in this position, a strong background in SRE practices, systems, networking, coding, capacity management, cloud operations, continuous delivery and deployment, and open-source cloud enabling technologies like Kubernetes and OpenStack is required. Deep understanding of the challenges and standard methodologies of running large-scale distributed systems in production, solving complex issues, automating repetitive tasks, and proactively identifying potential outages is also necessary. Furthermore, excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. As a Senior SRE at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now!

What you ll be doing:

  • Develop software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases.

  • Gain a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks.

  • Create tools and automation to reduce operational overhead and eliminate manual tasks.

  • Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation.

  • Define meaningful and actionable reliability metrics to track and improve system and service reliability.

  • Oversee capacity and performance management to facilitate infrastructure scaling across public and private clouds globally.

  • Build tools to improve our service observability for faster issue resolution.

  • Practice sustainable incident response and blameless postmortems

What we need to see:

  • Minimum of 5-8 years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments.

  • Master's or Bachelor's degree in Computer Science or Electrical Engineering or CE or equivalent experience.

  • Strong understanding of SRE principles, including error budgets, SLOs, and SLAs.

  • Proficiency in incident, change, and problem management processes.

  • Skilled in problem-solving, root cause analysis, and optimization.

  • Experience with streaming data infrastructure services, such as Kafka and Spark.

  • Expertise in building and operating large-scale observability platforms for monitoring and logging (e. g. , ELK, Prometheus).

  • Proficiency in programming languages such as Python, Go, Perl, or Ruby.

  • Hands-on experience with scaling distributed systems in public, private, or hybrid cloud environments.

  • Experience in deploying, supporting, and supervising services, platforms, and application stacks.

Ways to stand out from the crowd:

  • Experience operating large-scale distributed systems with strong SLAs.

  • Excellent coding skills in Python and Go and extensive experience in operating data platforms.

  • Knowledge of CI/CD systems, such as Jenkins and GitHub Actions.

  • Familiarity with Infrastructure as Code (IaC) methodologies and tools.

  • Excellent interpersonal skills for identifying and communicating data-driven insights.

NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.


Employment Type: Full Time, Permanent

Read full job description

Prepare for Senior Site Reliability Engineer roles with real interview advice

People are getting interviews at Nvidia through

(based on 55 Nvidia interviews)
Campus Placement
Job Portal
Company Website
Referral
Walkin
Recruitment Consultant
36%
25%
11%
7%
4%
4%
13% candidates got the interview through other sources.
High Confidence
?
High Confidence means the data is based on a large number of responses received from the candidates.

What people at Nvidia are saying

Senior Site Reliability Engineer salary at Nvidia

reported by 3 employees with 7-9 years exp.
₹28.8 L/yr - ₹97 L/yr
100% more than the average Senior Site Reliability Engineer Salary in India
View more details

What Nvidia employees are saying about work life

based on 512 employees
66%
96%
85%
79%
Flexible timing
Monday to Friday
No travel
Day Shift
View more insights

Nvidia Benefits

Free Transport
Free Food
Cafeteria
Health Insurance
Work From Home
Job Training +6 more
View more benefits

Compare Nvidia with

Qualcomm

3.8
Compare

Intel

4.2
Compare

Advanced Micro Devices

3.8
Compare

Texas Instruments

4.1
Compare

Broadcom

3.3
Compare

Applied Materials

3.9
Compare

Analog Devices

4.1
Compare

NXP Semiconductors

3.8
Compare

Sterlite Technologies

3.8
Compare

Indus Towers

3.9
Compare

Nokia Networks

4.3
Compare

Cisco

4.2
Compare

Lumen Technologies

4.0
Compare

Redington

4.0
Compare

Colt Technology Services

4.4
Compare

RadiSys

4.1
Compare

Vindhya Telelinks

4.1
Compare

Juniper Networks

4.2
Compare

ITI

3.7
Compare

Tejas Networks

4.1
Compare

Similar Jobs for you

Site Reliability Engineer at NVIDIA

Bangalore / Bengaluru

7-9 Yrs

₹ 32.5-37.5 LPA

Senior Site Reliability Engineer at NVIDIA

Hyderabad / Secunderabad, Pune + 2

7-9 Yrs

₹ 32.5-37.5 LPA

Senior Site Reliability Engineer at Barracuda Networks

Bangalore / Bengaluru

4-10 Yrs

₹ 25-30 LPA

AI Engineer at NVIDIA

Pune

5-7 Yrs

₹ 25-30 LPA

Site Reliability Engineer at InfraCloud Technologies

Mumbai

4-10 Yrs

₹ 25-30 LPA

Cloud Platform Engineer at Experian PLC

Hyderabad / Secunderabad

4-7 Yrs

₹ 22.5-30 LPA

Senior Architect at NVIDIA

Bangalore / Bengaluru

7-11 Yrs

₹ 20-27.5 LPA

Site Reliability Engineer 2 at PhonePe

Bangalore / Bengaluru

0-8 Yrs

₹ 30-35 LPA

Senior Site Reliability Engineer at Barracuda Networks

Bangalore / Bengaluru

4-10 Yrs

₹ 25-30 LPA

Senior Site Reliability Engineer at FabHotel Aay Kay Model Town

New Delhi, Karnal

7-10 Yrs

₹ 25-30 LPA

Nvidia Bangalore / Bengaluru Office Locations

View all
Bengaluru Office
NVIDIA Graphics PVT LTD, C-1 "Jacaranda", Wing-A Manyata Embassy Business Park, Outer Ring Road Bengaluru
Karnataka 560045
Bengaluru Office
Nvidia Graphics Pvt Ltd, C1, Nagavara Bengaluru
Karnataka 560045

Senior Site Reliability Engineer, Data Science and ML Platforms

5-8 Yrs

Hyderabad / Secunderabad, Pune, Gurgaon / Gurugram +1 more

2d ago·via naukri.com

Senior Software Engineer

4-7 Yrs

Bangalore / Bengaluru

2d ago·via naukri.com

Senior System Software Engineer

4-12 Yrs

Bangalore / Bengaluru

2d ago·via naukri.com

System Software Engineer, Conversational AI

1-5 Yrs

Pune, Bangalore / Bengaluru

2d ago·via naukri.com

Verification Engineer, CPU Performance Analysis

4-8 Yrs

Bangalore / Bengaluru

2d ago·via naukri.com

Senior Python Software Engineer, Security

4-8 Yrs

Bangalore / Bengaluru

2d ago·via naukri.com

Senior Site Reliability Engineer

7-9 Yrs

Hyderabad / Secunderabad, Pune, Gurgaon / Gurugram +1 more

2d ago·via naukri.com

Senior Responsible AI Engineer

5-7 Yrs

Pune

2d ago·via naukri.com

Senior Software QA Automation Engineer

5-7 Yrs

Bangalore / Bengaluru

2d ago·via naukri.com

Senior Network Engineer - Cloud

4-7 Yrs

Mumbai, Hyderabad / Secunderabad, Pune +2 more

2d ago·via naukri.com
write
Share an Interview