Loading jobs…
Loading jobs…
Metrostar Systems — Washington, District of Columbia
As a Sr. Systems Engineer II , you’ll deploy, operate, and support our observability platform — a containerized monitoring stack (TimescaleDB, OpenSearch, Telegraf, Promethus) running on Linux hosts in an on-premises, air-gapped environment — and to serve as a strong network troubleshooter for the broader environment. You'll keep data flowing end-to-end: from network telemetry collection, through ingestion and storage, to the dashboards stakeholders rely on.
A significant part of this role is diagnosing and resolving network problems independently, including unfamiliar issues across systems you didn't build or deploy. This is a role for someone equally comfortable reading a packet capture, debugging a systemd unit, and writing a CI pipeline. with the goal to make an impact across the federal government.
We know that you can’t have great technology services without amazing people. At MetroStar, we are obsessed with our people and have led a two-decade legacy of building the best and brightest teams. Because we know our future relies on our deep understanding and relentless focus on our people, we live by our mission: A passion for our people.
Value for our customers. If you think you can see yourself delivering our mission and pursuing our goals with us, then check out the job description below!
What You’Ll Do
: Deploy, configure, and maintain the monitoring stack across development and on-prem production environments. Lead network troubleshooting across the environment — isolate and resolve connectivity, latency, routing, and performance issues, often on systems and topologies you have no prior context for. Operate and troubleshoot Linux servers — services, storage, performance, logs, and security hardening.
Manage containerized workloads with Docker / Docker Compose; build and maintain images and service definitions. Manage software delivery into an air-gapped network: offline package/image mirroring, artifact transfer, and controlled update procedures. Diagnose data-flow issues across the network: DNS, routing, firewalls, TLS, load balancing, and switching/routing fabric.
Build and maintain CI/CD pipelines and infrastructure-as-code for repeatable, automated deployments. Write and maintain automation scripts (Bash, Python) for provisioning, monitoring, and routine operations. Manage time-series and search data stores (TimescaleDB/PostgreSQL, OpenSearch): backups, retention, indexing, and query performance.
Monitor system health, set up alerting, respond to incidents, and lead root-cause analysis. Maintain runbooks and documentation; participate in an on-call rotation as needed. What you’ll need to succeed: 8 years of experience Proficiency with network and system diagnostic tools such as Wireshark, tcpdump, dig, curl, traceroute, mtr, ss/netstat, and packet-level analysis.
Strong Linux administration experience (Ubuntu/Debian and RHEL), including shell scripting, systemd, package management, permissions, log analysis, and Bash automation. ), Infrastructure as Code (Terraform, Ansible), and Git-based workflows. Experience designing, deploying, and supporting on-premises, segmented, air-gapped, and disconnected environments, including local registries and package mirrors.
Skills in observability, automation, and security best practices, with experience using Grafana, Prometheus, OpenTelemetry, Python, SQL/PostgreSQL, secrets management, and SSH administration at scale. Knowledge of networking fundamentals, including TCP/IP, DNS, HTTP/HTTPS, TLS, routing, switching, subnetting, VLANs, and firewall configuration.
Qualifications
, skills, and relevant experience.