System Administration and Maintenance
Administration is now performed both with AI and for AI: triage, runbooks, and configuration are increasingly drafted by agents holding privileged tool access, while the fleet must also carry model-serving stacks under accelerator, cost, and latency budgets. The course keeps its operating-system and administrative-domain core as graded operational work and adds telemetry analytics, measured evaluation of AI-assisted operations, and accountability for machine-initiated changes.
Current description → proposed description
This course presents concepts and skills the professional system administrator must understand to effectively maintain enterprise information technology. Topics include: operating systems, application packages, administrative activities, and administrative domains.
This course presents the concepts and skills the professional system administrator must command to maintain enterprise information technology, including operating systems, application packages, administrative activities, and administrative domains. Graduate students work at the level of policy and architecture and are held to the operation itself: least-privilege identity, configuration as code, patch and release execution, backup and recovery proven by measured restore, capacity planning, and service-level commitments across heterogeneous fleets. The course also treats the contemporary reality that administration is increasingly performed with and for artificial intelligence. Students assemble observability pipelines that gather logs, metrics, and traces through vendor APIs and analyze them with statistical and machine-learning methods; evaluate AI-assisted incident triage and runbook generation against records of past incidents; and administer model-serving stacks as one more class of application package in the fleet, with accelerator capacity, inference endpoints, cost and latency budgets, and versioned rollback held to the same change-control discipline as the rest of the estate. Attention is given throughout to the risks automation introduces: prompt injection carried in logs and ticket text, over-permissioned agent tooling and package supply chains, unreviewed machine-generated configuration, and the privacy of operational data sent to third-party providers. Students defend guardrails, audit trails, and human escalation paths for privileged operations, and report incidents and residual risk responsibly to technical and non-technical stakeholders.
What changes
- Fleet operations graded as performed work: patch and release execution, validated restore, identity and directory operations under service-level objectives
- Least-privilege, revocable credentials and change-approval gates extended to service accounts and AI agents
- Observability pipeline built and analyzed with statistical and machine-learning methods
- Model-serving stacks administered as one more class of application package under the fleet's existing change control
- Measured evaluation of AI-assisted incident triage and runbook generation, with guardrails, audit trails, and human escalation for privileged operations
6 proposed outcomes, mapped to 6 program outcomes
Each outcome below is written to be observable and assessable, and each is mapped to the program learning outcomes for which it produces evidence.
Students will be able to implement the administrative lifecycle of a heterogeneous fleet of enterprise operating systems, application packages, and administrative domains, including provisioning and configuration as code, patch and release execution, backup and recovery proven by a measured restore against stated recovery-point and recovery-time objectives, identity and directory operations that issue scoped, revocable credentials under change-approval gates to human administrators, service accounts, and AI agents holding production tool access, and the API gateway and credential broker through which LLM-based agent and model services reach production, all held to stated service-level objectives.
PLO 6.1: among the components the student builds and operates is the API gateway and credential broker named in the CLO's final clause, which is the production path by which LLM-based agent and model services are served through APIs, and standing that path up is structured deployment of LLM services in an enterprise environment.
Students will be able to construct an observability pipeline that collects, normalizes, and stores logs, metrics, and traces from heterogeneous systems through vendor APIs and time-series stores, and analyzes that telemetry with statistical and machine-learning methods to detect incidents and forecast capacity.
PLO 2.1: the CLO requires collecting, cleaning, storing, and analyzing data from multiple sources using APIs, databases, and machine-learning methods, applied to the real operational problems of incident detection and capacity planning.
Students will be able to administer model-serving application packages within the enterprise fleet, including accelerator capacity allocation, serving endpoints behind an API gateway with quotas and rate limits, cost and latency budgets, and versioned rollout and rollback of model and prompt-template changes, applying the same provisioning, patch, and change-control discipline used for the rest of the estate.
PLO 6.1: operating served model endpoints through an API gateway with quotas, budgets, and versioned rollback is the structured deployment of LLMs in an enterprise environment, assessed here from the administrator's side as one class of application package in the fleet.
Students will be able to evaluate an AI-assisted incident triage and runbook-generation workflow against a held-out corpus of historical incidents, reporting classification precision and recall, time to mitigation, and characteristic failure modes, and refining retrieval sources, escalation thresholds, and tool permissions from the diagnostic results.
PLO 6.2: measuring precision, recall, and time to mitigation on held-out incidents and then revising the workflow from those diagnostics is the assessment of performance and limitations followed by iterative refinement that this outcome specifies.
Students will be able to justify a guardrail and escalation policy for automated administration that addresses prompt injection carried in logs and ticket text, over-permissioned or compromised agent tooling and package supply chains, unreviewed machine-generated configuration, and the export of operational data containing personal information to third-party model providers, specifying audit logging and human-in-the-loop approval for destructive operations.
PLO 1.1 and PLO 3.2 are stated word for word twice in the workbook, once under Christian Faith and once under Integrated Disciplinary Knowledge, and carry identical rows, so this CLO supports both equally: deciding whose personal data leaves the enterprise and who answers for a privileged change no person authored is exactly the privacy, accountability, and harm reasoning the outcome names.
Students will be able to communicate the cause, remediation, and residual risk of a production incident in a written postmortem and an oral briefing for non-technical stakeholders, stating plainly where automated or AI-assisted tooling contributed to detection, to diagnosis, or to the fault itself, and who remains accountable for the outcome.
PLO 5.1: the postmortem and the stakeholder briefing are graded artifacts that must convey technical findings and the ethical question of accountability for machine-initiated action clearly and responsibly to both technical and non-technical audiences.
Program outcomes this course reaches
Filled cells are program learning outcomes with at least one supporting course learning outcome in this course. Sparse coverage is expected — no single course carries all twelve.