Senior Engineer – AI Monitoring & Control Capabilities
Job Description:
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.
We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.
Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!
Job Description:
This is a critical-delivery role responsible for engineering and operating vendor applications end to end within an on-premises OpenShift platform, directly supporting the enterprise's AI monitoring, controls, and governance commitments. You will develop and integrate application components, deploy and scale containerized services, and provide the production support that keeps these services reliable, performant, and audit-ready.
The domain is AI Agent operations. The applications you support deliver guardrails and observability for LLM powered systems across the enterprise, enabling the model observability, traceability, and evaluation needed to meet risk management, compliance, and regulatory expectations. Because this capability underpins key delivery milestones, the role demands strong engineering discipline, sound operational judgment, and a sense of urgency.
The ideal engineer learns new tools and environments quickly, brings a pragmatic build-and-operate mindset, and can lead cross-functional initiatives without formal authority.
Responsibilities:
- Own delivery of critical vendor applications from build through production support in a fully on-premises OpenShift environment, meeting committed milestones and quality expectations.
- Build and maintain APIs and integrations connecting vendor guardrail and observability platforms to enterprise AI Agent solutions.
- Establish and manage workflows that monitor performance, safety, latency, throughput, quality, and guardrail effectiveness in support of AI risk and control objectives.
- Implement and maintain model evaluation frameworks using LLM-as-a-Judge and specialized language models.
- Develop observability dashboards and alerting using Grafana or similar tools, producing audit-friendly evidence and traceability.
- Troubleshoot production issues and improve scalability, resilience, and operational efficiency to sustain availability for critical business services.
- Contribute to evaluation methodologies and guardrail controls for agentic and multi-step AI systems.
- Lead and mentor engineers, establish development standards, and drive process improvements across cross-functional delivery.
- Provide input into financial planning, budget tracking, forecasting, and management of supported applications and services.
Required Qualifications:
- Bachelor's or Master's degree in Computer Science, Data Science, MIS, a related field, or equivalent practical experience.
- 13+ years of software engineering experience designing, building, and scaling production applications.
- Strong proficiency in Python, including common data and ML libraries. Hands-on experience deploying and operating containerized workloads on OpenShift or Kubernetes.
- Proven experience taking LLM- or RAG-based applications into production and maintaining them at scale.
- Experience building APIs and integrating third-party platforms.
- Working experience with AI Agents development, including OpenAI APIs and LangChain.
- Practical understanding of inference performance optimization. Experience applying guardrails, observability, and model-based evaluation to AI applications.
- Experience delivering within Agile and Agile-at-scale environments against firm timelines.
- Proven ability to influence outcomes and meet delivery commitments in a matrixed organization without direct authority. Strong problem-solving, communication, and stakeholder management skills.
Desired Qualifications:
- Experience delivering technology solutions within a large, complex, highly regulated enterprise, preferably financial services.
- Fast learner able to ramp quickly on new technologies and tools.
- Experience with annotation pipelines, feedback loops, fine-tuning, and model alignment techniques.
- Familiarity with prompt versioning and lifecycle management.
- Experience with vendor-provided guardrail and observability platforms such as Galileo, LangSmith, and Splunk.
- Working knowledge of additional AI Agents frameworks such as LlamaIndex or Haystack.
- Familiarity with evaluation and guardrails for agentic, multi-step AI workflows.
- Collaborative and self-directed work style supporting engineering teams.
Skills:
- Automation
- Influence
- Result Orientation
- Stakeholder Management
- Technical Strategy Development
- Application Development
- Architecture
- Business Acumen
- Risk Management
- Solution Design
- Agile Practices
- Analytical Thinking
- Collaboration
- Data Management
- Solution Delivery Process
Shift:
1st shift (United States of America)Hours Per Week:
40