Skip to main content

Incident Operations Specialist

Zapier
Remote - NAMERUpdated 55d ago
Compensation
$119k–$179k
Published range
Location
Remote - NAMER
Remote eligibility
Employment
Full-time
Mid-level
Role family
Operations
B2B SaaS
Apply on jobs.ashbyhq.com
Job actionsApply now
Job actionsApply now

About the job

About the role

As Zapier expands into the enterprise market and accelerates AI-driven development, incident management is increasingly critical to customer trust and operational reliability. The Incident Operations Specialist keeps that program running day to day through reliable tooling, clean data, repeatable workflows, and AI-powered automation. You’ll report to the Incident Program Manager and help shape how Zapier responds to incidents, learns from them, and supports the people doing that work. This is an operations role with technical depth, not a software engineering role.

What you'll do

  • Own incident tooling operations: keep incident.io, PagerDuty, on-call rotations, escalation paths, and Slack-based workflows configured, reliable, and integrated.
  • Build and maintain AI-powered workflows: thread summarisation, postmortem drafting, follow-up triage, severity classification, data hygiene.
  • Analyze incidents and drive improvement: participate in incidents and postmortems, spot recurring patterns, surface program-level friction with recommended fixes.
  • Operate data and reporting: build and troubleshoot dashboards and reports (Databricks, Grafana, Looker).
  • Sustain the IC community: grow the community of practice for Incident Commanders and Support Leads, coach responders, keep documentation usable under pressure.

What we're looking for

  • AI fluency (required, not optional): use AI-native tools (Cursor, Claude, Copilot, or similar) as default working environment; built AI-powered workflows that keep running when offline; can quantify how AI has changed throughput or quality.
  • Incident response and analysis experience: hands-on with incident.io and PagerDuty; on-call rotations, escalation paths, routing, integrations; analysed incidents after the fact and turned patterns into program-level improvements.
  • Technical depth: write SQL against Databricks, wire up API integrations, build Slack workflows, prototype lightweight AI agents.
  • Async-first and visible: status in public channels, close the loop, prioritise ruthlessly, translate technical detail into plain language for Support, GTM, and leadership.

Our stack

Incident: incident.io, PagerDuty, Slack. Data and observability: Databricks, Grafana, Looker, SQL, Datadog, Prometheus, Opensearch, Graylog. AI: Cursor, Zapier AI, Claude, or equivalent. Collaboration: GitLab, Coda, Google Workspace, Jira, Zendesk.

How to apply

Apply through the provided link. Answer a few in-depth questions in the application; resume and CV fields are optional. Education is not a requirement. Zapier is an equal-opportunity employer and will provide reasonable accommodations; contact jobs@zapier.com if needed.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.