
Keeping a library of medical articles current as guidance moves
Read the full post on AWS BlogFlo Health publishes thousands of medical articles a year on women's health, read by millions of people who are often looking things up at a bad moment. Medical guidance keeps changing underneath that library, and checking every article by hand as it does is slow. Monotonous work is also exactly where human reviewers start making mistakes.
MACROS is the system we built with them on Amazon Bedrock. It reviews existing articles against a set of medical guidelines, flags anything that has drifted out of date, and proposes revisions matching both current standards and Flo's editorial voice. A second process runs the other way, pulling new guidelines out of published research and preparing them for review. Content arrives as PDFs, plain text or Flo's own JSON, through a Streamlit interface on ECS or straight into S3 for scheduled runs, with Step Functions orchestrating the pipeline.
The proof of concept held 80% accuracy and over 90% recall in spotting content that needed updating, and cut per-guideline detection from hours to minutes. The more interesting finding was that the system applied guidelines more consistently than manual review did. Consistency, rather than raw accuracy, is where tired humans tend to lose to machines.
Two things turned out to be load-bearing: chunking long documents before assessment, and having medical experts validate the parsing rules rather than only the output. The conclusion the team kept reaching was that this works as augmentation, with experts staying in the loop for final validation.


