FLAT is a multimodal system that unifies image and text representations into a single sequence of continuous tokens, trained jointly for cross-modal alignment and generation tasks. It uses nested dropout to organize information hierarchically, enabling flexible inference trade-offs between computational cost and output detail through prefix selection.