Network Analyst I (Trainee)

Gruve · Pune, Maharashtra, India

Spotted 1h ago

What you'll need to apply

Fields this application requires

NameEmailRelocation answerNotice periodSalary expectations

Company-specific questions

  • This role requires to work in 24*7 rotational shift (Including Night Shifts), Are you comfortable for this?
  • What is your current CTC?
  • What is your current address?
  • I acknowledge my right to refrain from disclosing any confidential or proprietary information during the interview and understand that the company will not solicit such information, in accordance with the Non-Disclosure Agreement
  • I consent to the use of AI-powered note-taking tools during the interview process and acknowledge that the interview may be recorded and stored for evaluation purposes.
  • I hereby provide my consent to share my professional referees with Gruve and authorize Gruve, or its designated third-party representative, to contact them for the purpose of conducting business reference checks as part of the hiring process.
Job description

About this role

Employer-provided description, formatted for easier reading.

About Gruve

Gruve is an innovative software services startup dedicated to transforming enterprises to AI powerhouses. We specialize in cybersecurity, customer experience, cloud infrastructure, and advanced technologies such as Large Language Models (LLMs). Our mission is to assist our customers in their business strategies utilizing their data to make more intelligent decisions.

As a well-funded early-stage startup, Gruve offers a dynamic environment with strong customer and partner networks.

Position Summary

Entry-level engineer on the 24×7 NOC monitoring rotation.

First responder to alerts from the AI Fabrik network estate — data-center fabric, edge routers/firewalls and out-of-band console management — and to PulseAI infrastructure health alerts (GPU servers, control-plane/infrastructure nodes, OpenShift cluster nodes, switches and storage) arriving through the Gruve outbound collector, working strictly from runbooks under the shift senior and building toward independent first-level diagnostics within 6 months.

Key Roles & Responsibilities

  • Monitor the network monitoring dashboards and the PulseAI infrastructure dashboards (Grafana); acknowledge fabric, edge, firewall and infrastructure alerts within the tier SLA and perform first-level triage per documented runbooks.
  • Classify alerts (actionable / informational / false alarm) with evidence and record complete investigation notes in the ITSM ticketing platform (the SLA system of record).
  • Watch GPU-server, control-plane/infrastructure node, switch and storage health signals — device reachability, interface and link errors, optics, CPU/memory, GPU utilisation/temperature, OpenShift node status, storage capacity thresholds, telemetry reachability — log deviations and escalate per the severity matrix with accurate context and timelines.
  • Recognise PulseAI infrastructure severity conditions — node or GPU loss, front-end network unreachable, back-end RoCEv2 fabric degraded, out-of-band management unreachable, storage capacity critical — and escalate to the L1/L2 engineer on shift.
  • Run scheduled health checks on device reachability, SNMP/syslog/streaming-telemetry feeds and collector connectivity; raise telemetry-gap and monitoring-coverage tickets for newly onboarded equipment.
  • Assist with configuration backup verification, asset/inventory updates (firmware, BIOS, GPU driver levels) and topology record upkeep under supervision.
  • Execute clean shift handovers and maintain the shift log.
  • Complete the structured training path (routing and switching, EVPN-VXLAN fabric concepts, next-generation firewalls, Linux and GPU-server basics, and container / Kubernetes / OpenShift fundamentals including cluster networking) and participate in shift drills.

Must-Have Skills & Qualifications

  • BE/BTech (CS/IT/E&TC) or equivalent.
  • Fundamentals of networking (TCP/IP, subnetting, VLANs, routing and switching basics, DNS, DHCP) and operating systems (Linux, Windows).
  • Conceptual understanding of data-center network components — switches, routers, firewalls, out-of-band management — and of common fault types (link down, high utilisation, device unreachable).
  • Conceptual grasp of containers and Kubernetes/OpenShift (pods, nodes, namespaces, Services, Ingress) and of how cluster networking depends on the underlying fabric — enough to read platform and infrastructure health alerts and follow runbook-driven triage.
  • Clear written English for ticket and handover quality.
  • Committed to 24×7 rotational shifts including nights and weekends.

Good-to-Have Skills

  • Entry-level networking certification (e.g., CCNA, JNCIA or equivalent).
  • Home-lab exposure (Packet Tracer / GNS3 / EVE-NG); basic Python or shell scripting.
  • Public cloud networking fundamentals (any major provider).
  • Home-lab or foundation-level exposure to Kubernetes/OpenShift (KCNA, Red Hat OpenShift fundamentals) and basic Linux / GPU server concepts.
  • Familiarity with SNMP/syslog concepts and any monitoring tool (Grafana, Zabbix, LibreNMS or similar).

Why Gruve

At Gruve, we foster a culture of innovation, collaboration, and continuous learning. We are committed to building a diverse and inclusive workplace where everyone can thrive and contribute their best work. If you’re passionate about technology and eager to make an impact, we’d love to hear from you.

Gruve is an equal opportunity employer. We welcome applicants from all backgrounds and thank all who apply; however, only those selected for an interview will be contacted.

Interested in this role?Continue on Gruve's careers page.
Apply on Gruve