About the Job
NVIDIA is hiring for the Site Reliability Engineer role in Bengaluru. The position is focused on supporting SRE initiatives that improve the reliability, scalability, and efficiency of enterprise systems while giving candidates an opportunity to work with modern cloud-native technologies.
In this role, you will contribute to distributed systems that support NVIDIA's AI-powered enterprise products and services. You will also get hands-on exposure to database operations, infrastructure automation, observability, monitoring, incident response, and Kubernetes-based infrastructure.
The role involves working with technologies such as Python, TypeScript, JavaScript, Go, AWS, Azure, GCP, Docker, Kubernetes, Terraform, Linux, Git, PostgreSQL, Prometheus, Grafana, and OpenTelemetry. You will work alongside Cloud, Platform, Security, and AI/ML teams while learning established SRE and incident-management practices.
NVIDIA also encourages AI-assisted engineering practices, including coding agents and LLM-powered tools. Personal projects, internships, coursework, open-source contributions, hackathons, and hands-on experience with cloud infrastructure, automation, DevOps/SRE, or CI/CD can help candidates stand out.
Job Overview
|
Company
|
NVIDIA
|
|
Job Role
|
Site Reliability Engineer
|
|
Hiring Type
|
Fresher / Entry-Level Hiring
|
|
Job Type
|
Full-Time
|
|
Education
|
BS degree in Computer Science or a related technical field such as Physics or Mathematics, or equivalent practical experience
|
|
Location
|
Bengaluru, India
|
|
Work Model
|
Work From Office |
|
Job ID
|
JR2023532
|
|
|
Eligibility Criteria
01
Candidates should have a BS degree in Computer Science or a related technical field such as Physics or Mathematics, or equivalent practical experience.
02
Candidates should have foundational proficiency in at least one programming language such as Python, TypeScript, JavaScript, or Go.
03
Basic understanding of cloud platforms such as AWS, Azure, or GCP and containerization technologies including Docker and Kubernetes is expected.
04
Exposure to or coursework in Infrastructure-as-Code tools such as Terraform, AWS CDK, or CloudFormation is useful, along with a willingness to learn.
05
Familiarity with Linux/Unix systems, networking fundamentals, and Git is required.
06
Candidates should have strong problem-solving skills, curiosity, willingness to learn, and good communication and teamwork skills.
Key Responsibilities
01
Support SRE initiatives focused on improving the reliability, scalability, and developer efficiency of enterprise systems.
02
Assist in building and maintaining distributed systems that support NVIDIA's AI-powered enterprise products and services.
03
Help automate database operations, including provisioning, scaling, backup, and failover for relational and vector database services.
04
Contribute to observability and monitoring by building dashboards, alerts, and automation scripts to improve system performance and reliability.
05
Participate in incident response processes, including issue triage, reducing mean time to resolution (MTTR), and contributing to post-incident reviews.
06
Collaborate with Cloud, Platform, Security, and AI/ML teams to support platform reliability and implement SRE best practices.
07
Learn to operate and troubleshoot complex systems, including Kubernetes-based and cloud-native infrastructure, while following established system design and incident management standards.
08
Explore and adopt AI-assisted engineering practices, including coding agents and LLM-powered tooling, to improve day-to-day development workflows.
Skills Required & Skills to Add in Resume
Technical Skills
Key technologies and areas relevant to this role
PROGRAMMING
Python, TypeScript, JavaScript, Go
CLOUD
AWS, Azure, GCP, Cloud Infrastructure, Cloud-Native Development
CONTAINERS & IaC
Docker, Kubernetes, Terraform, AWS CDK, CloudFormation, Infrastructure-as-Code
SYSTEMS
Linux/Unix, Networking Fundamentals, System Troubleshooting, Distributed Systems
OBSERVABILITY
Logging, Metrics, Tracing, OpenTelemetry, Prometheus, Grafana, Monitoring
DATABASE
PostgreSQL, MySQL, SQL, Indexing, Query Optimization, Database Operations
DEVOPS & SRE
Site Reliability Engineering, DevOps, Automation, CI/CD, Incident Response, Reliability Engineering
VERSION CONTROL
Git, Collaborative Development Workflows
If you have actually used these technologies through projects, internships, coursework, or practical experience, highlight your strongest Python, Linux, Git, cloud, Docker, Kubernetes, Terraform, SQL, monitoring, automation, CI/CD, and SRE experience. Projects involving cloud deployment, infrastructure automation, observability, or troubleshooting can also strengthen your resume.
Why Apply?
This role can be a strong opportunity for candidates who want to build a career in Site Reliability Engineering, cloud infrastructure, and modern platform engineering. The position provides exposure to systems that support NVIDIA's AI-powered enterprise products and services.
You will have the opportunity to work with technologies such as Kubernetes, Docker, cloud platforms, Infrastructure-as-Code, databases, observability tools, and automation while learning how reliable distributed systems are operated at scale.
The role is also a good fit for candidates who enjoy troubleshooting, automation, and learning new technologies. Personal projects, internships, coursework, and hands-on experience can all be relevant when demonstrating your interest in SRE and cloud engineering.
Benefits
Health & Well-Being
Healthcare and programs supporting employee well-being.
Learning & Development
Learning resources and opportunities that support professional growth.
Financial Benefits
Financial programs and benefits available to eligible employees.
Time Off
Paid time-off and holiday programs, subject to applicable policies.
Benefits may vary based on location, employment status, and applicable company policies.
How to Apply
01
Visit the Official NVIDIA Careers Website
Click the Apply Now button below to open the official NVIDIA careers page for the Site Reliability Engineer position.
02
Open the Job Application
Review the Site Reliability Engineer job details and make sure your skills and qualifications match the requirements before starting your application.
03
Complete the Application
Enter your personal, educational, and professional information accurately as requested in the NVIDIA application form.
04
Upload Your Resume
Upload your latest resume and highlight relevant skills such as Python, cloud platforms, Docker, Kubernetes, Linux, Git, databases, automation, and observability.
05
Submit Your Application
Review the information you have provided and submit your application through the official NVIDIA careers portal. Keep an eye on your registered email for any further communication.
Disclaimer
The information shared in this job post is based on the job details available from the respective company at the time of publishing. Job requirements, eligibility, experience, work model, benefits, application process and other details may change at any time. Candidates are advised to verify the latest information on the company's official careers page before applying.
TechJobsAlert is not the hiring company and does not take part in the recruitment process. We do not guarantee selection, interview calls or employment. Candidates should apply only through the official company application portal or the application link provided in the job post.
Candidates are responsible for checking their eligibility and providing accurate information during the application process. Please refer to the official company website for the most up-to-date information, hiring guidelines and any instructions provided during the recruitment process.
Frequently Asked Questions
Is the NVIDIA Site Reliability Engineer role suitable for freshers?
The job description does not specify a minimum number of years of professional experience. It asks for a bachelor's degree or equivalent practical experience along with foundational technical knowledge, so candidates with relevant projects, internships, coursework, or practical experience may consider applying.
What educational qualification is required?
NVIDIA asks for a BS degree in Computer Science or a related technical field such as physics or mathematics, or equivalent practical experience.
What skills are required for this role?
The role requires foundational proficiency in Python, TypeScript, JavaScript, or Go, along with knowledge of cloud platforms, Docker, Kubernetes, Linux/Unix, networking, Git, relational databases, and observability concepts.
Is Kubernetes required for this job?
The job description asks for a basic understanding of containerisation technologies such as Docker and Kubernetes.
Which cloud platforms are relevant?
NVIDIA mentions AWS, Azure, and GCP. Candidates should have a basic understanding of at least one of these cloud platforms.
Can personal projects or internships help?
Yes. NVIDIA specifically mentions personal projects, internships, coursework, open-source contributions, hackathons, CI/CD experience, automation, and container orchestration as ways candidates can stand out.
Where can I apply for this NVIDIA job?
You can apply through the official NVIDIA careers portal using the Apply Now link provided in this job post.