
What this course assumes, the four boundaries every later section hangs off, and how to work the seven labs so they teach you something.
Why a working demo does not reach production. The five gating layers between a prototype and a deployed system, and which of them is yours.
Consultant, sales engineer, solutions architect, product engineer and FDE. Each has a defensible hand-off point, and confusing them causes most role disputes.
The FDE Trinity — engineer, strategist, implementer. Three personas held in one head, each with a signature move and a characteristic failure mode.
The four looping stages — diagnose, prototype, evaluate, adopt — each with an explicit failure mode, and how to read artifacts to tell which stage a project is really in.
Configure, integrate, extend or build — the four-way last-mile decision, and why picking 'build' by default is the most expensive mistake in the role.
Set up the toolchain and the engagement repository you will use for the rest of the course, including the seven evidence slots every engagement should fill.
Why nobody hands an FDE a specification, and why the requirements you are given describe a solution someone already chose.
Executives describe the workflow they believe exists. Operators describe the one that does. Why you interview the second group first.
Rewriting questions so they return evidence rather than opinion. 'Has this ever failed?' becomes 'show me the last rejected record'.
Two artifacts that carry an engagement: the assumption log, which records what you have not verified, and the interface catalogue, which records every way data moves.
The undocumented Friday CSV export that Finance depends on. Why shadow interfaces are invisible in every diagram and load-bearing in every workflow.
Reachability, identity and authorization are three separate failures with one symptom. A valid credential on an unreachable host looks exactly like a bad password.
Profiling the data you were promised against the data that exists. One country column holding UK, GB, GBR and EMEA is a semantic problem, not a cleaning problem.
Hands-on. Profile a deliberately messy customer extract, quantify what is actually wrong with it, and produce the readiness findings you would take back to the customer.
Turning discovery findings into a go/no-go the customer can disagree with, and why a scorecard with a blocked row is more useful than a green one.
The first deliverable is a written problem statement the customer agrees with. Why clarity precedes code, and what it buys you three months later.
What goes on the page: the problem in their words, the success metric, the explicit out-of-scope list, and the assumptions you are proceeding on.
The narrow path — one channel, one intent, one backend, fully instrumented and time-boxed. Why a horizontal slice produces plumbing and a narrow one produces evidence.
Choosing the slice. Three tempting wrong answers — the easiest path, the demo-polish path, and running two paths in parallel.
Acceptance across four dimensions — technical, business, adoption and operational — each with its own owner and its own sign-off.
Architecture decision records that survive a review six months later: the decision, the alternatives, the reason, and what would make you revisit it.
Hands-on. Take a set of raw, contradictory interview notes and produce a one-page brief with a success metric, an out-of-scope list and a named narrow path.
Transport success is not business correctness. A worked example where every request returned 200 and Finance still rejected the batch — and what the FDE should have checked instead.
System of record, authoritative source, field-level ownership and write authority are four different things. Getting them confused is how bidirectional syncs start oscillating.
Protocol-level idempotency and business-level idempotency are not the same guarantee. This lecture separates them, then introduces the dedup state machine the next lecture builds.
Hands-on. Three thousand orders through a flaky ERP, three client strategies, one seed. Watch a naive retrying client create 222 duplicate orders and an idempotency key remove all of them — using fewer API calls, not more.
Not every failure is retryable, and retrying the wrong ones multiplies. Permanent, transient, throttled and ambiguous — plus the arithmetic that turns three retries at five layers into 1,024 calls.
Offset, cursor and keyset pagination, and why a cursor without a durable checkpoint is not resumable. Includes the extractor that survives being killed halfway through a 400,000-row pull.
A webhook receiver is an admission-control boundary, not a function call. Duplicate delivery is normal, the sender retries on its own schedule, and doing real work inside the handler is how you get amplification.
Hands-on. Build a webhook receiver that admits fast, deduplicates durably, verifies the signature over raw bytes, and hands work to a worker — then hammer it with duplicate and out-of-order deliveries.
Partial failure, and the outcome you cannot classify. The payment processor authorised the charge and the response was lost — what your system is allowed to conclude.
The dual-write problem: you cannot atomically write to your database and publish a message. The transactional outbox, and exactly what it does and does not guarantee.
Hands-on. Build an outbox, kill the process between the business write and the publish, and show that no event is lost and no business record is duplicated.
Where duplicates re-enter a system, and why the retention window on a dedup table is a correctness property rather than a storage optimisation.
Timeout budgets and deadline propagation. A call with 120 milliseconds of budget left must not start a fresh 700-millisecond attempt.
Circuit breakers, bulkheads, load shedding and bounded queues. What must open a breaker, what must not, and why there is no fourth outcome under load.
Hands-on. Build a four-tier reconciliation — count, total, key-set, field-level — find the injected drift, and repair it through a queue with an immutable pre-action report.
Deploying into an environment you do not own, under a change process you did not design, with an approval you cannot skip.
Containerising the prototype, and the twelve-factor assumptions every platform quietly expects you to have met before it will run your image.
Rolling updates, readiness probes and rollback. What maxUnavailable actually does during a customer release window, and why a liveness probe is not a readiness probe.
Private networking as the thing that decides whether a deployment is permitted at all: VPC endpoints, PrivateLink, VNet integration, egress proxies and split-horizon DNS.
When the data genuinely cannot leave: self-hosted inference, what it costs you in operations, and the questions to ask before promising it.
Hands-on. Deploy the prototype behind a private endpoint with a readiness gate, then break the DNS resolution deliberately and diagnose it from the symptoms alone.
Why customer data is never tutorial data. Inconsistent formatting, bad OCR, duplicated documents, and permissions attached to every row.
The five stages of a retrieval pipeline — loading, indexing, storing, querying, evaluation — and the vocabulary to explain to a customer what you are building.
Chunking strategies and their failure modes: fixed-size, sentence-window, semantic and structural. Why the right answer depends on the document, not the framework.
Contextual retrieval — contextual embeddings plus contextual BM25 — and the measured reduction in retrieval failure it produces for an afternoon of work.
Hands-on. Build retrieval over a genuinely messy permissioned corpus, measure it against a twenty-question eval set, then improve it and prove the improvement.
Permissions inside retrieval rather than bolted on after. Why filtering results post-hoc leaks information, and how to scope the index instead.
When retrieval is the wrong answer: small stable corpora, questions requiring computation, and workflows where the real problem is that nobody wrote the document.
Workflow or agent — the decision that saves a month. Start with the simplest mechanism and add autonomy only when it demonstrably earns its place.
The five composable patterns — prompt chaining, routing, parallelisation, orchestrator-workers and evaluator-optimiser — with the conditions each is actually for.
Tool design as an interface problem: naming, parameters, error messages, and evaluating a tool rather than assuming the model will cope with it.
MCP and A2A as integration boundaries rather than model features, and what that reframing changes about how you secure and evaluate them.
The duplicate-order incident in detail: a generic execute-request tool, a timeout, a retry, and two orders in the customer's ERP.
Bounded autonomy in practice: approval gates, canonical request hashing, kill switches, and the metrics that tell you an agent is drifting.
Hands-on. Build a transactional write agent with an idempotency claim, an approval bound to a request hash, and a kill switch — then try to make it double-write.
Why demos lie. Selection effects, the questions you happened to try, and the difference between a system that works and one you have watched work.
Building the first eval set: twenty questions with expected answers, drawn from real usage, and why twenty honest cases beat two hundred generated ones.
Hands-on. Write a twenty-question eval set against your own corpus, score a baseline, make one change, and prove whether it helped.
LLM-as-judge: when it correlates with human judgment, when it does not, and the failure where the judge and the system share the same blind spot.
Evaluating agents on trajectory as well as outcome. A right answer reached by a wrong path is a system you cannot ship.
SLIs, SLOs, error budgets and the four golden signals — the vocabulary every conversation about 'is it reliable enough' actually runs on.
Hands-on. Instrument the pipeline with OpenTelemetry using the GenAI span conventions, then trace a slow request to the span that caused it.
Cost as a reliability concern: per-request budgets, the runaway retry loop that costs more than the outage, and what to show the customer monthly.
This course contains the use of artificial intelligence.
The demo went well. Six weeks later nothing is in production, the customer has gone quiet, and nobody can say exactly what went wrong.
Closing that gap is the job. A forward deployed engineer embeds with an enterprise customer, works out what the real problem is, builds inside their environment and their constraints, and stays until the thing is running and used. Part engineer, part product manager, part consultant. The role pays what it does because very few people are all three at once, in front of a customer, under time pressure.
This course teaches that work end to end. It is built around one project rather than a playlist of topics. In Section 3 you pick a messy document set you did not create — a mailbox export, a folder of scanned PDFs, a filings archive. You frame the problem, prototype against it, build an eval set, instrument it, take it through a security review, deploy it behind a private endpoint and hand it over with a runbook. Nothing restarts. By Section 11 you hold the two artifacts hiring managers actually screen for: a system that runs with monitoring and measurable quality, and a written brief an executive would read.
The technical half is current. Workflows against agents and the five composable patterns. Retrieval over data with real permissions attached, not a tutorial corpus. Evaluation as a discipline — offline and online, LLM-as-judge and where it lies to you, trajectory against outcome for agents. OpenTelemetry tracing with GenAI span conventions. SLOs and error budgets. Prompt injection, excessive agency, and the OWASP and NIST vocabulary a customer's security team will use on you. Containers, rolling updates, and the private-networking patterns that decide whether a deployment is allowed to exist.
The other half is the part engineers under-invest in and then lose deals over. How to run a discovery interview that surfaces what someone did last week instead of what they say they want. How to write a one-page brief with an explicit out-of-scope list — the section most people skip and the one that saves the engagement. How to demo the customer's before and after rather than your features. How to diagnose why a shipped system is not being used, and fix the actual barrier instead of running more training.
Some honesty about scope. This is not a course on training models; you will integrate and evaluate them, not build them from scratch. It assumes you can already write Python and call an API, and it skips the fundamentals to spend the time on the things that are hard to learn from documentation. If you have never deployed anything, start elsewhere and come back.
The course is twelve sections and eighty-seven lectures, and it moves in four movements. Sections 1 to 3 are discovery, scoping and deciding under constraint: how to interview an operator so they tell you what they did last week rather than what they want, the assumption log and interface catalogue, the shadow interfaces nobody documents, and the decision records that make a choice defensible six months later. Sections 4 to 6 are integration: HTTP 200 is not business correctness, idempotency at protocol and business level, the dual-write problem and the transactional outbox, timeout budgets, circuit breakers, and deploying into someone else's cloud behind a private endpoint.
Sections 7 to 9 are the part most people think is the whole job — retrieval over permissioned data, agents and tool design, evaluation and observability. Sections 10 to 12 are the part that decides whether any of it ships: the security review, prompt injection and blast radius, adoption, handover, and a capstone that carries one engagement from discovery to handoff.
Twelve of those lectures are hands-on, and each ships a downloadable bundle that runs on your machine with no cloud account: the code, the fixtures, a README, and a file showing the exact output a correct run produces. You will profile a deliberately messy customer extract and quantify what is actually wrong with it. Build an idempotent client with a dedup state machine and watch duplicates vanish. Run an outbox through a real SIGKILL mid-publish and prove nothing was lost or duplicated. Stand up retrieval over a corpus with permissions attached, then watch a naive retriever leak across a boundary and a filtered one not. Several of them break on purpose, because diagnosing the break is the lesson.
Three ideas recur and are worth naming, because they are what the course is really about. First, that a demo which works proves almost nothing — the interesting question is what happens at the ninetieth percentile, on the worst input, when a dependency is slow. Second, that authority is a design decision: what a system may do, and what it may do without a human, are two separate grants, and collapsing them is why security reviews refuse things wholesale. Third, that shipped is not adopted, and the barrier is usually something you built rather than something the users lack.
You will also leave with the two artifacts that get you hired. A system that runs, with traces and an eval suite and a number you measured before and after. And a written brief an executive would actually read. Section 12 is explicit about this: a screener spends four minutes deciding whether to book a call, and ten repositories give them nothing to hold on to, while a two-page engagement write-up with a measured outcome and six decision records naming the option you rejected gives them a sentence they can repeat.
The evaluation material deserves a specific mention, because it is where most courses wave. You will write an eval set before you tune anything — twenty cases drawn from real queries, cases the team says are hard, and cases that should refuse because no answer exists. Then you will watch a change improve the headline score and make the system worse: abstention lowers raw correctness and is kept anyway, because a user who cannot tell which ten per cent is wrong pays the checking cost on all of it. Ninety per cent accuracy with no confidence signal can deliver zero time saved, and that arithmetic is the reason adoption fails more often than latency is.
A word on who this is not for. If you have never shipped anything to production, the course will make sense and will not change what you can do — the material assumes you have felt the difference between code that works and code that survives a Tuesday. If you want to train models, this is the wrong course; you will integrate and evaluate them here, never build them. And if you are looking for framework tutorials, the frameworks change faster than any course can track, which is why this one teaches the constraints underneath them instead.
On tooling and dependencies: the course teaches patterns rather than products. You will meet OpenTelemetry, OWASP and NIST vocabulary, containers and private networking, because those are what a customer's security team and platform team will speak. But nothing here is tied to a vendor you must buy, and every lab runs against local stand-ins so the material outlives whichever provider you are using this year.
Every lab runs on a free tier. Every figure quoted in a lecture comes from code that ships with the course, and where a result was inconvenient it is reported as measured.