Collaborate with data scientists and AI engineers to:
* * Design, construction and operation of ML infrastructure * *
- Design and operation of ML infrastructure using AWS, GCP, Azure, etc.
- Create and optimize Docker images of Python code and ML models
- Container orchestration for production environments using Kubernetes (EKS/GKE/AKS)
- Infrastructure configuration management with Terraform/CloudFormation
- Cost estimation and optimization
* * Construction and operation of MLOps pipeline * *
- Design and operate CI/CD pipelines (GitHub Actions, etc.)
- Automate your ML pipeline with Airflow, Kubeflow, Vertex AI/SageMaker Pipelines, and more
- Experiment management and model registry operation using MLflow etc.
* * Model deployment and monitoring * *
- Model API deployment using KServe, SageMaker Endpoint, etc.
- System and model monitoring with Prometheus, Grafana, Datadog, etc.
-Monitoring, logging and alarm design of model performance (latency, accuracy, etc.)
- Automated retraining cycles and recommendations for improvement
* * Required skills and experience * *
Any of the following:- More than 3 years of development experience in languages such as Python
- Experience operating cloud platforms such as AWS/GCP/Azure
- Hands-on experience with Docker, Kubernetes (EKS/GKE/AKS)
- Experience with IaC tools such as Terraform/CloudFormation
- CI/CD build and operational experience such as GitHub Actions
- Experience in development and operation in a Linux environment
- Fundamentals of machine learning or data analysis
- Cost optimization experience
- Japanese proficiency at business level and above (Reference standard: Japanese proficiency test N2 and above)