When i was playing with Dall-E 3, which before being "disconnected" was the best in my opinion for testing visual concepts, i learned some keywords and did the rest by myself, patiently step by step, so it began naturally easier to translate or test things i wanted to do.
at some point, i was asking to myself it the generator(s) understood what it was doing (i know, there a re a lot of debate related to that point) or if it was just mimicking things it had been trained on (spoiler alert : this is a mix of the two concepts), so i started to train it to create skeletons made of tiny houses instead of bones (i say "train it" because everything you sent as inputs were considered as new statistical possibilities, that's the main idea, for future creations), and little by little it did it, it was all about the understanding of the structure used in a picture.
The best way to think a generator is to think it as a 3D software in terms of scene building and lightings placements. It helps a lot