How COG and Zarr Compare and Where They Meet
Is Zarr the New COG? Betteridge’s law of headlines strikes again. The answer is no. The real point is that Jarrett Keifer uses the question to make a stronger case about how the formats compare and where they can still connect.
I don’t think we need a winner. Not everything has to be Zarr, nor COG. I think having multiple effective technologies is a much bigger win for the community than picking just one. “Is Zarr the new COG?” is catchy, but I think it’s the wrong question. “What are the strengths and weaknesses of Zarr and COG and how should I pick one or the other for a new data product?” just doesn’t have the same ring to it, I know, but that’s probably the better question.
The better question, yes. Keifer’s comparison is strongest when it shows not just how COG and Zarr differ, but what future interoperability between them can mean for the technology and community.
COG and Zarr are so deeply different that a shared stack between them is impossible. That belief is self-fulfilling: nobody funds a bridge they’re convinced can’t be built, and then the absence of the bridge gets cited as proof of the chasm. Except we keep watching the bridge get built the moment someone has a reason to pay for it. AWS needed NITF and JPEG 2000 readable as virtual Zarr, so the codecs got written and shipped in osml-imagery-io. virtual-tiff needed predictor 2, so now Zarr reads it. The missing shared codec layer isn’t evidence of a thick boundary between these formats; it’s evidence of a coordination failure. The winner-take-all framing is part of the failure, because money that believes it must pick a winner doesn’t fund the seam between them. […]
The places where these formats meet—shared low-level readers, shared codec implementations, conventions with tooling behind them—are where a dollar helps both ecosystems at once. obstore, async-tiff, VirtualiZarr, even GDAL at times: these all show the veneer can thin in different ways. Let’s capitalize on the good ideas these demonstrate.