Why doesn't Docker Hub cache Automated Build Repositories as the images are being built?

Question

Note : It appears the premise of my question is no longer valid since the new Docker Hub appears to support caching . I haven't personally tested this. See the new answer below .

Docker Hub's Automated Build Repositories don't seem to cache images. As it is building, it removes all intermediate containers. Is this the way it was intended to work or am I doing something wrong? It would be really nice to not have to rebuild everything for every small change. I thought that was supposed to be one of the best advantages of docker and it seems weird that their builder doesn't use it. So why doesn't it cache images?

UPDATE: I've started using Codeship to build my app and then run remote commands on my DigitalOcean server to copy the built files and run the docker build command. I'm still not sure why Docker Hub doesn't cache.

Answer 1

Disclaimer: I am a lead software engineer at Quay.io, a private Docker container registry, so this is an educated guess based on the same problem we faced in our own build system implementation.

Given my experience with Dockerfile build systems, I would suspect that the Docker Hub does not support caching because of the way caching is implemented in the Docker Engine. Caching for Docker builds operates by comparing the commands to be run against the existing layers found in memory.

For example, if the Dockerfile has the form:

FROM somebaseimage
RUN somecommand
ADD somefile somefile

Then the Docker build code will:

Check to see if an image matching somebaseimage exists
Check if there is a local image with the command RUN somecommand whose parent is the previous image
Check if there is a local image with the command ADD somefile somefile + a hashing of the contents of somefile (to make sure it is invalidated when somefile changes), whose parent is the previous image

If any of the above steps match, then that command will be skipped in the Dockerfile build process, with the cached image itself being used instead. However, the one key issue with this process is that it requires the cached images to be present on the build machine , in order to find and verify the matches. Having all of everyone's images on build nodes would be highly inefficient, making this a harder problem to solve.

At Quay.io, we solved the caching problem by creating a variation of the Docker caching code that could precompute these commands/hashes and then ask our registry for the cached layers, downloading them to the machine only after we had found the most efficient caching set. This required significant data model changes in our registry code.

If you'd like more information, we gave a technical overview into how we do so in this talk: https://youtu.be/anfmeB_JzB0?list=PLlh6TqkU8kg8Ld0Zu1aRWATiqBkxseZ9g

Answer 2

The new Docker Hub came out with a new Automated Build system that supports Build Caching.

https://blog.docker.com/2018/12/the-new-docker-hub/

Why doesn't Docker Hub cache Automated Build Repositories as the images are being built?

Question

2 answers

solution1
25 ACCPTED 2015-07-31 20:27:22

solution2
3 2018-12-27 13:47:17

Why doesn't Docker Hub cache Automated Build Repositories as the images are being built?

Question

2 answers

solution1 25 ACCPTED 2015-07-31 20:27:22

solution2 3 2018-12-27 13:47:17

solution1
25 ACCPTED 2015-07-31 20:27:22

solution2
3 2018-12-27 13:47:17