Jobs / United States / Databricks

Senior Staff Applied AI Engineer - Context Retrieval

Databricks · 🇺🇸 Mountain View, California; San Francisco, California

Sponsorship verdict

Sponsorship possible

One solid signal, not two — worth applying, and worth asking about sponsorship early.

  • Employer is on a government sponsor recordThe US Department of Labor certified 445 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 78 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).
  • The posting doesn’t mention sponsorshipSilence isn’t a refusal — ask the recruiter before investing much time.
  • No salary bar for this routeH-1B has no fixed salary bar: the employer must pay at least the prevailing wage for the role and area. Cap-subject employers enter a lottery weighted by wage level. Source: https://www.federalregister.gov/documents/2025/12/29/2025-23853/weighted-selection-process-for-registrants-and-petitioners-seeking-to-file-cap-subject-h-1b, rules effective 2026-02-27.
  • What Databricks paid sponsored hires in similar roles103 certified filings for “Software Engineer” (Software Developers) in CA: $158k–$191k, median $188k. Most were filed at wage level II (47%) — 2 lottery entries, ≈31% projected selection odds for cap-subject employers. Source: US Department of Labor LCA disclosure data (Oct 2025 – Jun 2026).
  • Confirmed live todayWhen a source last listed this job as open.

US H-1B: cap-subject employers enter a lottery weighted by wage level — Level I gets 1 entry, Level IV gets 4 (DHS projected selection odds ≈15% at Level I to ≈61% at Level IV). Universities and non-profit research employers are cap-exempt. The $100,000 fee for new petitions from abroad is currently blocked by a court order (appeal pending).

A verdict summarises public evidence; it is not legal advice and never a guarantee — the employer and the immigration authority decide. Sign in to factor in where you can already work.

Start free →

Or apply yourself on the official page →

Sponsor Radar — Databricks

445 H-1B filings certified since Oct 2025

The US Department of Labor certified 445 H-1B/E-3 labor condition applications for this employer between Oct 2025 and Jun 2026 (latest Jun 2026) — the step every H-1B hire needs first. USCIS also records 78 H-1B approvals in FY2023. Source: LCA disclosure data (US Department of Labor (OFLC)).

Past sponsorship or register membership never guarantees sponsorship for this vacancy or for you. Full Sponsor Radar for Databricks →

About the role

P-1549 At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. The Mission Databricks agents are only as good as the context they can retrieve. Whether an agent is answering a question about last quarter's revenue, debugging a failing job, generating SQL against a 10,000-table lakehouse, or summarizing a Wiki page, its quality is bounded by what it can find — and how well it understands what it finds. We are hiring a Senior Staff Applied AI Engineer to own context retrieval for Databricks agents across SaaS providers . This is a zero-to-one role with two deeply connected charters: • Build the retrieval stack — query understanding, content understanding, ranking, retrieval, and evaluation — across the Enterprise SaaS data stored across multiple systems. • Build the search subagents that sit on top of that stack and reason about what context is needed , how to retrieve it , and whether the right thing actually came back — closing the loop between an agent's intent and the substrate that serves it. If you have deep Information Retrieval wisdom, have shipped retrieval systems for RAG and agentic workloads, and want to build the substrate — and the agents on top of it — that make every Databricks agent measurably smarter, this role is for you. What You Will Do • Build the full retrieval stack from scratch. Own the end-to-end system: query understanding, content understanding and indexing, hybrid retrieval, ranking, and evaluation. Make the architectural calls that will define how Databricks agents access context for years to come. • Retrieve across heterogeneous data — structured and unstructured. Index and rank across structured assets (tables, columns, SQL queries, dashboards, code, notebooks, jobs) and unstructured content (docs, wikis, tickets, chat, images, video, audio). Each modality has its own signals — design retrieval that exploits them rather than flattens them. • Connect to the SaaS surface area customers actually use. Build connectors and retrieval adapters for the systems where enterprise knowledge lives. Treat each retrieval source with its own freshness, permissions, and ranking signals. • Optimize for two consumers at once. Retrieval must serve both LLMs (grounded, token-efficient, hallucination-resistant context) and humans (intuitive, explainable discovery). These are different objectives and require different signals — own both. • Crack query understanding for agents. Agent queries don't look like web queries. Build query rewriting, decomposition, intent classification, and entity resolution tuned for multi-turn agentic workflows. • Crack content understanding at scale. Build the pipelines that extract structure, entities, embeddings, summaries, and metadata from every supported asset type — and keep them fresh as customer data evolves. • Build search subagents that reason about retrieval. Design the agentic layer that decides what context is needed , which sources to query , how to decompose and route the search , and — critically — whether the retrieved content is actually sufficient to answer the question . These subagents will plan multi-hop searches, issue follow-up queries when results are weak, ground claims against retrieved evidence, and hand back high-confidence context (or signal failure) to upstream agents. This is where IR meets agentic reasoning. • Build the evaluation flywheel for both retrieval and subagents. Stand up offline evals (nDCG, MRR, Recall@K, Precision@K), LLM-as-

View the official posting →

Source: Greenhouse (employer board) First seen: 2026-08-18 Last confirmed: 2026-10-03 How our data works → Report this job

Similar opportunities