Interview Process: 2 virtual rounds; onsite interview may also be requested
Note to Vendor to find a right candidate for this job:
Please find and share strong W2 candidates for this Production Support Engineer role in Chandler, AZ (3 days onsite/2 days remote). Candidates must have 3 5+ years of hands-on production support experience with strong Autosys, Unix/Linux, Shell scripting, SQL/Oracle, Splunk, Dynatrace, and ServiceNow. Strong troubleshooting skills are critical, including Unix commands, SQL performance/issues, Autosys job states, incident management (P1 P4), and real-time production incident/root-cause analysis. Please submit candidates who can confidently explain actual production support scenarios during a technical interview, not just candidates with matching resume keywords.
Job details
Job Title: Production Support Engineer
Job location: Chandler, AZ (3 Days onsite, 2 Days Remote) Hybrid role.
Job Type: 12+ Months contract
Job Summary
Team and role context
- This is a live incident production support seat, not a pure runbook follow role. Team monitors batch, feeds, and application health across Unix, SQL, Dynatrace, Splunk, and ServiceNow, and owns incidents end to end including root cause.
- Candidates need to be comfortable being handed a vague scenario, for example an application down two hours after a change went in, and walking an interviewer through live triage, not reciting a memorized process.
Required
- 3 to 5 years hands on application production support
- Strong Autosys
- Strong Unix and Shell scripting
- Working SQL, Oracle, Hadoop or similar DBMS
- Ability to juggle and prioritize multiple concurrent issues
- Fast independent learner
Required, confirmed from team member
- Command level Unix fluency, not tool familiarity. Interviewers ask for the actual syntax: du and df for disk space, top and uptime for system load, find with -type and -mtime flags for locating files by age. A candidate who can describe what a command does but can't produce the flags will get caught here.
- Ability to distinguish failure types precisely. Interviewers specifically probe whether a candidate conflates a data issue with a Unix file system issue with a database connection pool issue. These are three different diagnostic paths and the interviewer will correct and re-ask if a candidate blurs them.
- SQL depth beyond basic querying: index behavior and why a query runs slow, truncate versus delete, join types used in production troubleshooting, not just writing selects.
- Autosys job state knowledge beyond scheduling, specifically the difference between a job marked inactive versus on hold versus on ice, and why a team would use each.
- Splunk and Dynatrace used together, and candidates should be able to explain why a team runs both rather than just one. Splunk gets tested as a log search and correlation tool, not described as monitoring.
- Noisy alert management. Interviewers ask how a candidate would handle an alert threshold that's firing too often, for example tuning a CPU alert from 70 percent to 85 percent to cut false positives.
- Incident severity fluency: candidates should be able to define P1 through P4 without hesitation and describe how communication changes at each level, including whether they'd post to a status page versus direct email during an outage.
- Basic API and auth troubleshooting exposure, specifically JWT or token related login failures, came up as a possible scenario topic.
- Shell scripting for automation of repetitive work, for example disk cleanup or log rotation scripts, should be a real example the candidate has done, not something on the resume without a story behind it.
Interview process, confirmed
- Round 1 is scenario based and purely technical. No resume discussion. Interviewer gives a live incident scenario and expects the candidate to narrate triage, tools used, and resolution in real time.
- Round 2 is resume based. Interviewers read line by line and will ask candidates to defend anything listed on the skill matrix. Do not submit a resume with skills listed that the candidate can't speak to in detail.
Screening guidance for recruiters
- Resume keyword matches on Autosys, Splunk, ServiceNow, Unix are not sufficient signal for this req. Confirm during your screen that the candidate can talk through an actual production incident end to end, including specific commands and specific database failure types, not general process language.
- If a candidate's Unix experience is light, mostly GUI tools or managed environments, flag it before submission rather than letting round 1 be the first test.