Cerner (Oracle Health)
Program Manager 4
Skills
Job description
The IC3 TPM owns defined GPU cluster health workstreams and drives day-to-day execution for assigned customers, shapes, locations, repair programs, or partner actions. This role tracks availability, repair workflow adherence, operational KPIs, risks, issues, dependencies, and program milestones. The role is accountable for crisp execution, reliable status, early risk identification, and disciplined escalation.
Customer and Cluster Availability Tracking
Own availability tracking for assigned customer clusters defined by shape/location.
Monitor performance against the 97.5% availability target and identify clusters trending below target.
Maintain weekly views of unavailable hosts, long-running repairs, reopen rates, repair aging, spare constraints, and repair throughput.
Track repair progress through standard repair workflow states, including CPV/ Warminator , TRS/CPV Tier 1, DO/CHS, release, termination, and customer handoff.
Program Execution
Manage assigned repair improvement programs with clear scope, milestones, owners, risks, and status.
Drive execution of targeted programs such as proactive unavailable-host reduction, repair workflow adherence, RMA/spares follow-up, and partner action tracking.
Coordinate with SDE/SRE teams to ensure tooling, triage, and first-level support gaps are visible and prioritized.
Escalate blockers where repair SLAs, availability goals, or customer commitments are at risk.
KPI and Reporting Support
Support definition and tracking of business and operational KPIs, including availability, billability, revenue impact, repair success, repair efficacy, reopen rates, new tickets, and throughput.
Maintain dashboards, weekly reports, status summaries, and action registers.
Partner with engineering and data teams to improve reports, dashboards, forecasts, and tooling.
Provide structured input into leadership, customer, and internal team updates.
Stakeholder Coordination
Facilitate working sessions across TPMs, SDE/SRE, partner teams, and repair execution teams.
Document decisions, risks, open actions, and follow-ups.
Ensure stakeholders understand current repair status, next steps, and escalation path.