Find your next role
Discover amazing opportunities across our network of companies committed to gender equality in the workplace.
Software Engineering, Data Science
Taipei City, Taiwan
We are seeking an AI System Engineer to support hands-on implementation, validation, monitoring, and day-to-day technical operations for a large-scale AI Data Center platform. The role combines NVIDIA GPU infrastructure, virtualization, Kubernetes, monitoring, basic network and storage integration, and structured customer-facing technical support. The engineer will execute approved designs, test cases, change procedures, and runbooks, and will work with senior architects and platform specialists when issues require deeper escalation.
The role is suitable for an infrastructure engineer with practical GPU, VMware, Kubernetes, monitoring, or cloud-native deployment experience who is ready to expand into AI data center platform operations. Equivalent experience with comparable technologies may be considered for selected platform tools.
AI / GPU Infrastructure Deployment and Operations
· Support installation, configuration, validation, and operational checks for NVIDIA GPU servers and GPU-enabled infrastructure.
· Perform GPU node health checks, diagnostic collection, basic topology validation, and performance or functional testing.
· Support NVIDIA vGPU or physical GPU environments, including driver and runtime validation under approved procedures.
· Assist with server, OOB/BMC, operating system, certificate, DNS, NTP, firewall, and end-to-end connectivity checks.
· Collect system, GPU, platform, and application logs and prepare diagnostic evidence for escalation to senior engineers or vendors.
Virtualization, Kubernetes, and Platform Operations
· Support VMware-based infrastructure and cloud-native platform deployment, testing, and post-deployment troubleshooting.
· Deploy and maintain Kubernetes workloads using YAML and Helm, including namespaces, resource limits, services, and basic GPU scheduling.
· Support cluster onboarding, tenant or project setup, user access, quotas, and resource pools through Rafay or comparable Kubernetes management platforms.
· Execute approved provisioning, upgrade, retry, recovery, rollback, and configuration-validation procedures.
· Assist with platform integration testing across infrastructure, Kubernetes, monitoring, and customer applications.
Network and Storage Integration
· Perform first-level TCP/IP, VLAN, routing, DNS, NTP, firewall, and service-connectivity validation using documented procedures.
· Support approved tenant network, IP addressing, VLAN/VRF, route, and segmentation changes through Netris or comparable network automation tools.
· Assist with storage client installation, mounts, CSI integration, NFS/S3 access, and capacity or latency checks using Weka or comparable storage platforms.
· Validate tenant isolation, network reachability, storage connectivity, and change results; escalate advanced fabric or storage issues when required.
Monitoring, Troubleshooting, Incidents, and Changes
· Implement and operate monitoring agents, exporters, dashboards, and log collection using Prometheus, Grafana, or comparable tools.
· Perform routine GPU, system, Kubernetes, network, and storage health checks and identify abnormal conditions.
· Conduct first-level alert triage, problem reproduction, log analysis, ticket updates, troubleshooting, and escalation.
· Execute approved standard changes, updates, service restarts, certificate renewals, validation, and rollback during scheduled maintenance windows.
· Support incident, request, change, tenant onboarding, offboarding, and decommissioning activities with complete operational records.
Testing, Documentation, and Customer Support
· Execute functional, integration, regression, basic performance, recovery, and end-to-end test cases.
· Maintain configuration records, test evidence, incident notes, tickets, SOPs, runbooks, and known-error documentation.
· Prepare clear technical findings, status updates, and handover materials for internal teams, customers, and partners.
Support technical walkthroughs, operational guidance, and user or partner enablement sessions when required.
· Bachelor’s degree or higher in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
· Approximately 2-5 years of hands-on experience in system engineering, infrastructure support, virtualization, cloud-native platforms, data center operations, or a related technical role.
· Hands-on experience with NVIDIA GPU or vGPU infrastructure, GPU-enabled servers, or comparable accelerator environments.
· Practical experience with Kubernetes and containerized infrastructure; working knowledge of Helm and YAML is expected.
· Experience with VMware virtualization or a comparable enterprise virtualization platform.
· Experience using monitoring and troubleshooting tools such as Prometheus, Grafana, or comparable observability platforms.
· Basic Linux administration capability and the ability to interpret operating system, platform, and application logs.
· Working knowledge of TCP/IP and common infrastructure services such as VLANs, routing, DNS, NTP, and firewalls.
· Ability to follow runbooks, test cases, change tickets, and maintenance procedures accurately and to document results clearly.
· Strong customer-facing communication and technical coordination skills in project-driven environments.
· Full professional proficiency in Chinese for daily communication, issue handling, ticket updates, and operational documentation.
· Basic to intermediate English proficiency for technical manuals, product interfaces, vendor documentation, and support cases.
· Ability to work primarily on-site in Taiwan and support scheduled maintenance windows and occasional after-hours activities.
· Experience managing NVIDIA driver, firmware, CUDA, NCCL, GPU telemetry, or GPU performance benchmarking.
· Experience with Rafay, Netris, Weka, or comparable Kubernetes management, network automation, IPAM/SDN, or high-performance storage platforms.
· Exposure to NVLink, InfiniBand, Spectrum-X, RoCE, high-speed Ethernet, or other AI cluster interconnect technologies.
· Experience with enterprise storage, CSI, NFS, S3, backup/restore, snapshots, or storage performance monitoring.
· Experience supporting production environments using incident, problem, change, or ITSM processes, including NOC/SRE-style operations.
· Scripting or automation experience using Python, Shell, Ansible, APIs, or similar tools.
· Relevant certifications such as CKA, CKAD, NVIDIA AI Infrastructure/Operations, Linux, Kubernetes, networking, or storage certifications.
In a world where technology never stands still, we understand that, dedication to our clients success, innovation that matters, and trust and personal responsibility in all our relationships, lives in what we do as IBMers as we strive to be the catalyst that makes the world work better.
Being an IBMer means you’ll be able to learn and develop yourself and your career, you’ll be encouraged to be courageous and experiment everyday, all whilst having continuous trust and support in an environment where everyone can thrive whatever their personal or professional background.
Our IBMers are growth minded, always staying curious, open to feedback and learning new information and skills to constantly transform themselves and our company. They are trusted to provide on-going feedback to help other IBMers grow, as well as collaborate with colleagues keeping in mind a team focused approach to include different perspectives to drive exceptional outcomes for our customers. The courage our IBMers have to make critical decisions everyday is essential to IBM becoming the catalyst for progress, always embracing challenges with resources they have to hand, a can-do attitude and always striving for an outcome focused approach within everything that they do.
Are you ready to be an IBMer?
IBM’s greatest invention is the IBMer. We believe that through the application of intelligence, reason and science, we can improve business, society and the human condition, bringing the power of an open hybrid cloud and AI strategy to life for our clients and partners around the world.
Restlessly reinventing since 1911, we are not only one of the largest corporate organizations in the world, we’re also one of the biggest technology and consulting employers, with many of the Fortune 500 companies relying on the IBM Cloud to run their business.
At IBM, we pride ourselves on being an early adopter of artificial intelligence, quantum computing and blockchain. Now it’s time for you to join us on our journey to being a responsible technology innovator and a force for good in the world.
IBM is proud to be an equal-opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, gender, gender identity or expression, sexual orientation, national origin, genetics, pregnancy, disability, neurodivergence, age, or other characteristics protected by the applicable law. IBM is also committed to compliance with all fair employment practices regarding citizenship and immigration status.
For additional information about location requirements, please discuss with the recruiter following submission of your application.