Senior Platform Security Engineer
AI FactoryOS Operations
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.
Senior Platform Security Engineer Role Summary
Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Platform Security Engineer owns the operational security of that platform . The multi-tenant Kubernetes estate sits at its heart: production access, admission policy, workload identity and network policy in the live estate, and the operational security controls that keep it defensible as it grows .
This is a hands-on senior role with deep technical expertise . This role owns security in production: runtime visibility and enforcement built with modern eBPF -based tooling, the security acceptance requirements a release has to meet before it is authorised , the authority to grant or refuse a security exception, and the deepest technical escalation for security-related platform faults.
Key Responsibilities
- Build runtime visibility and enforcement using eBPF -based tooling (for example Cilium, Falco or Tetragon), feeding high-quality signal into the estate's security monitoring, and tune detection so that alerts stay actionable as the estate grows.
- Own the security gates in the delivery pipeline: image signing, software bills of materials, vulnerability gating and policy checks, and maintain the policy-as-code controls that govern what ships to production, running on the delivery toolchain operated by Shared Services Operations.
- Set and own the security content of the multi-tenant Kubernetes estate: admission policy, workload identity and network policy, and verify that what runs in production matches it.
- Set the operational security standard the platform's Kubernetes control planes and baselines must meet in production, validate against it continuously, and raise deviations as engineering requirements to the AI Infrastructure team.
- Validate the layered controls that limit the blast radius of a fault or compromise across tenant boundaries, test them against realistic failure and attack paths, and feed the findings to AI Infrastructure where the control belongs to the product.
- Investigate and respond to security-related operational anomalies, distinguishing a security signal from an ordinary operational symptom.
- Operate the platform's security controls in line with ISO 27001 and NIST frameworks, applying deep expertise in securing large-scale GPU infrastructure , and produce the security evidence those controls generate for collation by the Service Delivery Manager .
- Provide the deepest technical expertise for security-related platform faults, contribute to post-incident review for security incidents, and mentor engineers across the function on secure design and runtime security practice .
- Provide independent monitoring of privileged activity across the estate, including on the identity, certificate and secrets platforms, using the access mechanisms built by the Principal Platform Identity Engineer.
Skills Experience
Required Skills
- Strong experience of privileged access security, including least-privilege design, secrets management, and the auditing of privileged operations in production environments.
- Strong experience securing multi-tenant Kubernetes environments, including admission policy, workload identity, network policy and control plane hardening.
- Extensive experience with eBPF -based runtime security tooling (for example Cilium, Falco or Tetragon), including building detection and enforcement that feeds a security monitoring platform.
- Strong DevSecOps experience owning security gates in a delivery pipeline: image signing, software bills of materials, vulnerability gating and policy checks.
- Solid understanding of hardware-enforced or network-based tenant isolation, and how it interacts with Kubernetes-level policy.
- Experience with policy-as-code tooling for Kubernetes and delivery pipelines.
- Demonstrated experience as a senior escalation point for security-related faults in a 24/7 production environment, including post-incident review.
- Strong scripting or programming ability for security automation and tooling (for example Python, Go or Bash).
- Clear technical judgement and communication, with the ability to translate security findings and standards for engineering teams and leadership.
Preferred Experience
- Experience securing GPU or HPC infrastructure and multi-tenant AI workloads.
- Experience in a multi-tenant service provider, cloud or colocation environment .
- Experience with workload identity and secrets management patterns, and how runtime enforcement ties back to them.
- Experience with security frameworks and evidence production under ISO 27001, SOC 2 or NIST frameworks.
- Relevant security certification (for example OSCP, GCFA, or a Kubernetes security credential).
- A Bachelor's degree in computer science , engineering or a related discipline, or an equivalent combination of relevant experience and training.
Expected Outcomes
- Runtime detection covering the estate, with an alert set that stays actionable as the estate grows.
- Every production release meeting the security acceptance requirements, with exceptions owned, time-bound, compensating-controlled and on the register.
- The operational security standard for Kubernetes validated continuously in production rather than asserted at design time.
- Blast radius controls tested against realistic failure and attack paths, with findings closed or evidenced to the platform engineering teams.
- Security evidence produced through normal operation rather than assembled before an audit.
Location Reporting
Location : Based in Australia or Singapore, with travel to Australian AI Factory sites as required .
On-call : The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function.
Sourced from a public career listing. Jobverse is an aggregator, not the employer.