Senior Software Engineer, Amazon AGI · Researcher

Rohith Nama

I work on how large-scale AI systems get built, evaluated, and governed. My day job is production inference infrastructure at Amazon. My research looks at agentic AI benchmarking, literacy, and standards.

12 yrs software engineering Amazon AGI · Nova & Mantle ACM TechBrief co-author IEEE Senior Member
Rohith Nama
01 · About

About

I'm a Senior Software Development Engineer at Amazon, in the Artificial General Intelligence (AGI) organization, where I work on inference infrastructure and responsible-AI systems for the Nova family of foundation models. Before that, I worked on data infrastructure at Capital One.

Alongside engineering, I research how agentic AI systems should be evaluated. These are systems that plan and carry out multi-step tasks on a user's behalf instead of producing output for a person to review, and they bring oversight problems the field is still working out. Some of that work has reached policy channels, including a technology policy brief written for Congress and federal agencies and a public comment on NIST's draft benchmark evaluation standards.

02 · Engineering

Systems Work

AMAZON
NOVA / MANTLE
Inline multimodal responsible-AI validation
Designed infrastructure for Amazon's Nova and Mantle inference systems that runs safety, copyright, and provenance checks inline during generation, before output reaches a user. Doing this for multimodal systems is the hard part: the checks have to cover text, image, video, and speech inside a single inference pass without slowing it down.
AMAZON
NOVA 2 OMNI
GroupedContent Abstraction
Designed the architecture behind Nova 2 Omni's unified multimodal inference pipeline. It preserves the relationships between independent content streams, such as speech aligned to its transcript or an edited image aligned to its source, within one pipeline, and it underlies the model's multimodal capabilities at launch. AWS has publicly described Nova 2 Omni as its first unified multimodal reasoning model, in production use at enterprises including SAP, Deloitte, Palantir, Siemens, and Zendesk.
RESEARCH
KDD 2026 · ARXIV
SOP-Bench
A benchmark for evaluating LLM agents on real industrial standard operating procedures: over 2,000 tasks across 12 business domains, each with an executable tool interface and a ground-truth outcome. Accepted at KDD 2026 and released as open source. Research groups with no connection to the project have since cited it, including teams at Nanyang Technological University and a Meituan, Georgia Tech, and Nanjing University collaboration.
RESEARCH
GOVERNANCE
Agentic literacy debt
Originated the concept of agentic literacy debt: the deficit that accumulates when agentic systems are deployed at scale faster than users gain the capacity to understand, supervise, and contest their actions. The debt framing matters because the gap compounds. Each delegation normalizes the next, agents increasingly interact with other agents, and organizations that skip literacy infrastructure once rarely build it later. The AI Literacy Institute has covered the argument in its review of the field.
RESEARCH
AI AND ETHICS
(SPRINGER)
Agentic AI Literacy Framework (AALF)
Co-developed the framework that answers the literacy debt problem. AALF defines the competencies people need once they delegate decisions to AI agents instead of reviewing AI output: six competency domains across three proficiency levels, grounded in principal-agent theory. Published in AI and Ethics.
03 · Policy

Policy & Publications

ACM TECHBRIEF
SPRING 2026 · NO. 18
Agentic AI: Autonomy, Opportunities, and Challenges of Action-Taking AI Systems
Co-authored with Dr. Larry Medsker and Brian Peretti for the ACM Technology Policy Council. ACM TechBriefs are short, non-lobbying technical briefings sent to Congress, federal agencies, and NIST. Authors are invited by the committee on the basis of subject-matter expertise; there is no self-nomination.
PUBLIC COMMENT
NIST AI 800-2
Response to NIST's draft benchmark-evaluation practices
Contributed technical input to ACM's public comment on NIST's draft guidance for automated benchmark evaluations of language models and AI agents. The comment argues for scaffold-sensitivity analysis as standard practice in agent evaluation, for treating open-weight and hosted models differently in safety benchmarking, and for closing the multimodal evaluation gap, since text-only testing misses cross-modal safety risks in models that are deployed multimodally.
PEER REVIEW
& EDITORIAL
Reviewer, Springer Nature journals
Peer reviewer for Artificial Intelligence Review, The Journal of Supercomputing, and AI and Ethics. Also serving as Lead Guest Editor for the AI and Ethics topical collection Where Ethics Meets Engineering: Innovation-Led Approaches to AI Safety and Accountability.
04 · Speaking & Mentorship

Talks & Mentorship

TALK
CIOs, AI and the Outlook for IT Modernization
George Mason University, Center for Excellence in Government Cybersecurity Risk Management and Resilience.
TALK
AI and Data Analytics for Anti-Corruption Compliance Forum
The American Conference.
PANEL
Compliance in the Age of AI 2026
Opal Group. Featured main-stage panelist.
MENTORSHIP
GENER8TOR
Expert mentor and featured speaker, gener8tor
Mentored founders and spoke to a generative-AI cohort at the accelerator.
MENTORSHIP
TECHSTARS
Mentor, Techstars
Advises early-stage founders on applied AI and engineering.
MENTORSHIP
AWS
Mentor, AWS Impact Accelerator
Mentors founders building on AWS through Amazon's startup accelerator program.
05 · Recognition

Recognition

IEEE
Senior Member
Institute of Electrical and Electronics Engineers.
PRESS
FRAUD MAGAZINE
"Agentic AI is reshaping the future of fraud risk"
Feature interview for the Association of Certified Fraud Examiners' Fraud Magazine on agentic AI, inference-time safety controls, and fraud prevention.
PRESS
ACM SIGAI
AI MATTERS
AI ethics & policy column interview
Interviewed for ACM SIGAI's AI Matters on evaluation, literacy, and governance challenges specific to agentic AI systems.