Production backend
Topify.ai
From customer conversations to a production backend that runs the whole marketing workflow.
Python · FastAPI · ReAct workflows · BGE-M3 + BM25 + RRF + reranking · Railway
Start with the customer's workflow
B2B marketing work spans information gathering, analysis, and repeated actions across tools. Adding a model call does not resolve that fragmentation. The product challenge was to translate real SEO and generative engine optimization (GEO) workflows into software that customers could run repeatedly without an operator stitching the steps together.
I worked directly with customers to scope requirements, then owned the production Python/FastAPI backend and its deployment. The unit of delivery was one automated workflow: analyze a customer's AI-search visibility, retrieve the relevant material, execute the LLM and tool steps, generate content, and publish it to WordPress or another CMS.
System architecture
Retrieval combined BGE-M3 embeddings for semantic similarity, BM25 for exact terminology, reciprocal rank fusion to merge the two rankings, and a reranking stage before results reached the LLM workflow. The corpus exceeded 100,000 chunks.
Execution ran as multi-step ReAct workflows over more than 30 external APIs. Independent LLM queries within a step were issued in parallel, and the service was deployed on Railway as part of my end-to-end delivery responsibility.
Runtime behavior for long-running jobs
Parallelize what is independent
A single customer job fans out into many model and API calls. Issuing independent LLM queries concurrently kept job duration bounded by the slowest dependency chain rather than the sum of every call.
Bound every external call
Each API and model call carried a timeout, and per-provider rate limits were enforced inside the workflow so that one busy integration could not stall the whole job or exhaust a quota.
Fail partially, not totally
Partial-failure controls let a job continue when one of its 30+ integrations failed, recording what was missing instead of discarding the completed steps. That is what made long-running customer jobs operable rather than fragile.
Delivery evidence & constraints
The retrieval latency figure is scoped to the retrieval step in this implementation. It is not an end-to-end LLM response time, a p95 measurement, or an uptime commitment.
The public summary does not include customer records, commercial source code, revenue figures, or confidential product materials. The Live link points to the current public product site. This case study describes my work during November 2025–January 2026; it does not attribute subsequent product changes to me.
What I bring to the next system
This work demonstrates customer-facing requirements discovery, production backend ownership, concurrency and failure handling across many external dependencies, retrieval implementation, and deployment. It is the part of my background where the problem definition and the running service lived in the same role.
I carry that connection into infrastructure work: understand the operating need, define the workflow, and follow it through to the running system.