How AI makes and reads images

You will be able to explain how image generators and image-reading models work and what each is bad at.

Joanne runs a small bakery in Tampines. On Monday she asks an image generator for a poster for her Hari Raya cookie promotion, and in seconds she has a warm, glossy photo of cookie jars on a festive table. It looks professional until she reads the banner in the picture, which says "Hari Raua Speical", and notices that one jar has a lid floating beside it. On Wednesday she photographs a crumpled supplier invoice and asks an assistant to add up the flour orders. It gives a neat total, a little over what she expected.

Both tools work with images, but they do opposite jobs. One makes pictures from words, and the other reads pictures and answers in words. This lesson covers how each one works and where each tends to slip.

Making a picture from noise

Many image generators use a method called diffusion. The idea is strange at first but easy to follow.

During training, the builders take millions of images with captions and add random noise to them step by step, like static on an old television, until nothing of the original is left. The model is trained on the reverse job: given a noisy image and its caption, predict what noise to remove to get one step closer to a clean picture. Lesson 2.2 described training as guess, measure the miss, adjust. Here the guess is "this is the noise", and the miss is measured against the noise that was really added.

To make a new image, the generator starts from pure random noise. It removes a little noise, guided by your text prompt, then a little more, over many steps. A picture gradually forms. Your prompt steers each step towards an image that fits the words, which is why "cookie jars on a festive table" produces jars and a table rather than a beach.

There is no stored photo being copied or edited. The generator is building something new that matches the patterns it learned from its training images. Not every generator uses diffusion, and builders combine methods, but the point holds for all of them. The image is assembled from patterns the model learned, and at no stage does it work from a plan of what should be where.

Reading a picture

Models that read images work differently. Lesson 3.1 showed that words are cut into tokens and turned into numbers before the model sees them. Images go through a similar process. The picture is cut into small patches, and each patch is turned into a list of numbers the model can work with, much as a word token is. Those image tokens go into the context window alongside your question.

That is why you can upload a photo of a whiteboard, a chart from a report or a screenshot of an error message and ask about it in plain words. The model is predicting text, as always, but now the text in front of it includes the picture in a form it can use.

What image generators get wrong

Generators are good at the general look of a picture, such as its lighting, style and mood. They are weaker at anything that has to be exactly right.

Text inside images is the classic problem, as Joanne's banner showed. The model learned what writing on a banner tends to look like, more than the precise spelling of the words. Counts are another: ask for seven cupcakes and you may get six or nine. Hands, with their fingers and joints, have long been a weak spot. So has consistency. If you need the same mascot in five different poses for a campaign, the face, colours and outfit tend to drift from one image to the next.

These tools keep improving, and some of these weaknesses are much less common than they were. But the safe habit is to check every generated image for text, counts, hands and small details before anyone outside your team sees it.

What image readers get wrong

Image readers are useful for pulling text and figures out of photos and screenshots. The risk is that they misread, and do so with confidence.

Trouble comes from small print and crumpled receipts, from blurred photos and handwriting, and above all from dense tables. A 6 can become an 8, a column can shift by one row, and a smudged figure can be replaced with one that looks plausible. When a value is hard to read, the model does not leave a gap. It predicts the likeliest text, which may be a number that was never on the page. Joanne's total was a little high because one quantity had been read as 18 instead of 10.

So treat any figure an image reader extracts as a draft. Check it against the original, especially totals, dates, amounts in S$ and anything you will pay, claim or report. Lesson 6.3's checking routine applies here just as it does to text answers.

The clearest way to see this is to give a reader something with lots of numbers and compare line by line. A receipt from your last grocery run or a chart from a work report will do.

Ask an assistant to read a photo of a receipt or a chart and list every number it extracted, then mark which ones it got wrong.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).