The Motern Method
A book I wanted to read wasn’t on Kindle. So I filmed the pages on my phone, and my agent turned the video into a website I can read on my phone.
Where it started
There’s a book I like called The Motern Method, by Matt Farley. One guy who says to make all your ideas. He’s published tens of thousands of songs, for example.
But it’s not on Kindle, and I read exclusively on the app. I asked my agent (My setup) to find a digital copy. There isn’t one anywhere, so I bought the physical book.
The last chapter is the behind-the-scenes bit, how he made the book using the lessons in the book. I got most of the way there and the kids came back from school. I wanted those last pages on my phone so I could dip in and out.
I thought, I can just film myself flicking through the pages and send that to my agent. So I did.
I had some trouble uploading the video from my phone (think it was too big a file), so I AirDropped it to my Mac and pointed my agent at where it was, in my Downloads.
analyse the video frames and extract the text and put into one html file and publish with here.now - video is in ~/Downloads/IMG_7131.MOV
What I’ve got, what I want back, and where it should end up. One html file, put on here.now so it has a link.
Reading the pages
It took longer than I thought it would. First it cut the video into frames.
Then it had to read them. It tried OCR first, which wasn’t good enough, then a model with vision, which did work. This is the first frame, both ways.
Some of the frames were too big to send, so it shrank them and tried again. Then it wrote the page. Eight times. Each rewrite drifted from what it had read, so it went back and rewrote it again.
How it works
I was testing an open-source model, DeepSeek, for this task, so I was curious how much it cost. About 21 cents, plus about 15 cents for the vision model. Most of that was the agent re-reading its own session: as a session grows, each new turn re-reads all the context.
Checking it
Anything it couldn’t read, under my thumb or in the curve of the page, is in pale italics. I should’ve said to not include the partial first page but it’s an agent, it read my instructions and did what it was told.
It told me the text was word for word. Here’s page 126.
The first two paragraphs are gone, and the rest is in its own words. The page is in the video, clear as anything. Next time I’d ask it to check every page against the frame it came from (Talking to agents) before it publishes.
Make it your own
Anything in the wrong shape for how you read works the same way. A recipe in a screenshot, the handout from a talk, the whiteboard at the end of a meeting, a letter you need to reply to. Film it or photograph it, and tell your agent what you’ve got, what you want back, and where it should end up.
i filmed the handout from today’s talk, it’s in my downloads. pull the text out into one html file i can read on my phone. mark anything you can’t read, then check every page word for word against the video before you publish it with here.now
Notes
- The book wasn’t on Kindle, so I filmed myself flicking through the pages and sent the video to my agent.
- What I’ve got, what I want back, and where it should end up.
- It cut the video into frames. OCR wasn’t good enough, a model with vision was.
- About 36 cents, most of it the agent re-reading its own session.
- It said word for word. Page 126 wasn’t. Next time I’d have it check every page against its frame first.



