To Apply for this Job Click Here
As an AI Software Engineer, you’ll join a small, senior delivery team building and operating a direct-to-teacher AI tool backend and the AI platform around it. We’re looking for someone who can pick up well-scoped features and own them through to production with light support – someone who has shipped production AI systems before and understands how they behave once real users are in the loop.A note on shape: this is a full-stack-leaning-backend role with a strong Al-systems emphasis. It is not a pure machine-learning role and not a pure web role. Most of your time is in async Python and the LLM orchestration layer; you’ll touch the frontend when the work calls for it.
- EST hours are preferred.
- LinkedIn Profiles are required on the resume.
Must Have for this role:
- 5+ years of experience in software development with a strong focus on backend engineering.
- Ideal candidate has hands-on experience with AI pipelines and workflow development, Pinecone vector databases, AI agents and agent orchestration, AI guardrails and governance concepts.
- Experience with LangGraph and LangSmith is a strong advantage.
- Candidate should be able to contribute immediately with minimal ramp-up time.
- Open to hiring a strong Full Stack Engineer, but the role has a heavier backend focus.
- Python is the primary programming language used by the team.
- Experience with FastAPI is highly desirable.
- Knowledge of Go (Golang) is a significant plus, as some components are being rewritten in Go.
- Familiarity with Azure and AWS cloud platforms is important.
- Understanding of modern AI application architecture and adaptability to evolving AI technologies is valued.
- Some with Front end experience in TypeScript, Angular.
What you’ll do:
- Build and extend the agentic LLM orchestration – the graph of nodes and agents that turns a teacher’s request into a useful response.
- Integrate external data sources and tools into the agent so it can reason over the information teachers need.
- Improve retrieval quality across vector and lexical
- Work on model routing, resilience, and graceful degradation so the system stays fast and available under real-world load.
- Strengthen the prompt lifecycle, evaluation discipline, and observability that keep a non-deterministic system reliable.
- Harden quality with automated unit, integration, and end-to-end testing.
What this role is – and isn’t
This is a software engineering role on a production AI system. You won’t be training or fine-tuning models or running model-science experiments – our data science team owns that. But unlike a generic application role, prompt engineering, retrieval quality, and evaluation discipline are core to this job, not someone else’s problem. The interesting work lives in the orchestration, the prompt lifecycle, retrieval, and evals – not CRUD.
What the role looks like at this level
As a Software Engineer II on this team, you’ll break down medium-sized features, estimate them, and cut scope to ship on time. You’ll start to own tasks within the service with support from senior teammates, contribute to technical design and engineering-review proposals while thinking through failure cases, give helpful and timely code reviews, and defend your decisions in review. You’ll debug to root cause in your area, instrument your code for operations, and participate in the on-call rotation. Senior engineers are around to pair with and review your work
- but increasingly you’ll be the one proposing the approach and carrying a feature to
What you must already bring
You don’t need every line below at expert depth, but the combined surface has to be covered.
Core engineering
Expert-level async Python (3.11+). Real production asyncio / async / await experience across the request path – a synchronous-only Python background won’t be enough here.
- FastAPI at depth: routers, dependencies, lifespan, middleware. Pydantic v2 and disciplined type hints.
- pytest and pytest-asyncio – fixtures, async, mocking, and meaningful coverage. Standard formatting, linting, and type-checking tools are table
- Hands-on production experience with LangGraph: state machines, conditional edges, checkpointing. Experience with LangChain alone is not the same thing – this is where most of the surface area lives.
- LangChain core (messages, runnables, tools), and prompt engineering/ prompt lifecycle management – versioned, environment-tagged prompts with local overrides – using tracing and experiment tooling such as LangSmith.
- Multi-agent/ multi-node workflow design – routing across specialized agents and
- RAG with hybrid vector + lexical retrieval, and experience with a managed LLM provider such as Azure OpenAI (deployments, API versions, quotas).
- Sound instincts for non-determinism, token budgets, timeouts, and graceful degradation, plus familiarity with eval frameworks (e.g. LLM-as-judge and regression evals).
Data, infrastructure, and delivery
- PostgreSQL operationally – indexing, connection pools, poolers – plus pgvector and OpenSearch/Elasticsearch hybrid hybrid (text + KNN) search.
- WS and Kubernetes in production – genuine fluency, beyond local container orchestration. Docker multi-stage builds; infrastructure-as-code (e.g. Terraform) and manifest overlays for multiple environments.
- Multi-environment configuration discipline – several environments, from local through production, each with its own secrets, prompts, and resources.
And comfortable with
- Typescript and modern Angular with RxJS when frontend work is needed. A backend-leaning candidate is welcome as long as you’re comfortable in Angular; a frontend-leaning candidate must still be solid in the Python/LLM
Nice to have {genuine bonuses, none required)
- MCP (Model Context Protocol) and SSE; database migration tooling; Redis-compatible
- Observability tooling (APM, metrics, tracing) and distributed-tracing
- Modern Python packaging and build tooling, Make-based builds, GitHub Actions, private package registries, and encrypted-secrets workflows.
- Load testing and end-to-end browser testing
- Edtech / K-12 domain awareness (standards, proficiency, learning frameworks) and FERPA-adjacent data-privacy thinking.
- Familiarity with large-enterprise internal identity, auth, and content-metadata services – accelerates ramp, but learnable.
How we work
This is an internal enterprise codebase, so expect internal SDKs and package registries, encrypted-secrets tooling, and a secrets manager as part of the daily flow. It’s a polyglot repo backend, frontend, infrastructure-as-code, and database migrations coexist – and the team uses written design and decision docs. Security hygiene for AI apps (prompt injection, PII handling, guardrails) matters here because we’re working with educational data.
Reference: 1061108
Worried that you don’t meet every single requirement listed in the job ad? Studies have shown that individuals from marginalized groups are less likely to apply to jobs unless they meet every single qualification. Hive + Co. is dedicated to building a diverse, inclusive and representative workplace, so if you’re excited about this role, but worried that you don’t meet every requirement, we encourage you to apply anyways. We’d love to get to know you.
