I've committed the initial modifications to the storage rebuilding code. The changes mostly live in the AUFS and COSS code - the rest of Lusca isn't affected.
The change pushes the rebuild logic itself into external helpers which simply stream swaplog entries to the main process. Lusca doesn't care how the swaplog entries are generated.
The external helper method is big boost for AUFS. Each storedir creates a single rebuild helper process which can block on disk IO without blocking anything else. The original code in Squid will do a little disk IO work at a time - which almost always involved blocking the process until said disk IO completed.
The main motivation of this work was the removal of a lot of really horrible, twisty code and further modularisation of the codebase. The speedups to the rebuild process are a nice side-effect. The next big improvement will be sorting out how the swap logs are written. Fixing that will be key to allowing enormous caches to properly function without log rotation potentially destroying the proxy service.
Tuesday, July 28, 2009
Monday, July 13, 2009
Caching Windows Updates
There are two issues with caching windows updates in squid/lusca:
* the requests for data themselves are all range requests, which means the content is never cached in Squid/Lusca;
* the responses contain validation information (eg ETags) but the object is -always- returned regardless of whether the validators match or not.
This feels a lot like Google Maps who did the same thing with revalidation. Grr.
I'm not sure why Microsoft (and Google!) did this with their web services. I'll see if I can find someone inside Microsoft who can answer questions about the Windows Update related stuff to see if it is intentional (and document why) or whether it is an oversight which they would be interested in fixing.
In any case, I'm going to fix it for the handful of commercial supported customers which I have here.
Wednesday, July 8, 2009
Storage rebuilding / logging project - proposal
I've put forward a basic proposal to the fledgling Lusca community to get funding to fix up the storage logging and rebuilding code.
Right now the storage logging (ie, "swap.state" logging) is done using synchronous IO and this starts to lag Lusca if there is a lot of disk file additions/deletions. It also takes a -long- time to rotate the store swap log (which effectively halts the proxy whilst the logs are rotated) and an even longer time to rebuild the cache index at startup.
I've braindumped the proposal here - http://code.google.com/p/lusca-cache/wiki/ProjectStoreRebuildChanges .
Now, the good news is that I've implemented the rebuild helper programs and the results are -fantastic-. UFS cache dirs will still take forever to rebuild if the logfile doesn't exist or is corrupt but the helper programs speed this up by a factor of "LOTS". It also parallelises correctly - if you have 15 disks and you aren't hitting CPU/bus/controller limits, all the cache dirs will rebuild at full speed in parallel.
Rebuilding from the log files takes seconds rather than minutes.
Finally, I've sketched out how to solve the COSS startup/rebuild times and the even better news is that fixing the AUFS rebuild code will give me about 90% of what I need to fix COSS.
The bad news is that integrating this into the Lusca codebase and fixing up the rebuild process to take advantage of this parallelism is going to take 4 to 6 weeks of solid work. I'm looking for help from the community (and other interested parties) who would like to see this work go in. I have plenty of testers but nothing to help -coding- along and I unfortunately have to focus on projects that provide me with some revenue.
Please contact me if you're able to help with either coding or funding for this.
Friday, June 26, 2009
Lusca in Production, #2
Lusca-head pushing >100mbit in production..

Here's the Lusca HEAD install under FreeBSD-7.2 + TPROXY patch. This is a basic configuration with minimal customisation. There's ~ 600gig of data in the cache and there's around 5TB of total disk storage.
You can see when I turned on half of the users, then all of the users. I think there's now around 10,000 active users sitting behind this single Lusca server.
Tuesday, June 23, 2009
Lusca-head in production with tproxy!
I've deployed Lusca head (the very latest head revision) in production for a client who is using a patched FreeBSD-7.2-STABLE to implement full transparency.
I'm using the patches and ipfw config available at http://tproxy.no-ip.org/ .
The latest Lusca fixes some of the method_t related crashes due to some work done last year in Squid-2.HEAD. It seems quite stable now. The bugs only get tickled with invalid requests - so they show up in production but not with local testing. Hm, I need to "extend" my local testing to include generating a wide variety of errors.
Getting back on track, I've also helped another Lusca user deploy full transparency using the TPROXY4 support in the latest Linux kernel (I believe under Debian-unstable?) He helped me iron out some of the bugs which I've just not seen in my local testing. The important ones (method_t in particular) have been fixed; he's been filing Lusca issues in the google code tracker so I don't forget them. Ah, if all users were as helpful. :)
Anyway. Its nice to see Lusca in production. My customer should be turning it on for their entire satellite link (somewhere between 50 and 100mbit I think) in the next couple of days. I believe the other user has enabled it for 5000 odd users. I'll be asking them both for some statistics to publish once the cache has filled and has been tuned.
Stay tuned for example configurations and tutorials covering how this all works. :)
I'm using the patches and ipfw config available at http://tproxy.no-ip.org/ .
The latest Lusca fixes some of the method_t related crashes due to some work done last year in Squid-2.HEAD. It seems quite stable now. The bugs only get tickled with invalid requests - so they show up in production but not with local testing. Hm, I need to "extend" my local testing to include generating a wide variety of errors.
Getting back on track, I've also helped another Lusca user deploy full transparency using the TPROXY4 support in the latest Linux kernel (I believe under Debian-unstable?) He helped me iron out some of the bugs which I've just not seen in my local testing. The important ones (method_t in particular) have been fixed; he's been filing Lusca issues in the google code tracker so I don't forget them. Ah, if all users were as helpful. :)
Anyway. Its nice to see Lusca in production. My customer should be turning it on for their entire satellite link (somewhere between 50 and 100mbit I think) in the next couple of days. I believe the other user has enabled it for 5000 odd users. I'll be asking them both for some statistics to publish once the cache has filled and has been tuned.
Stay tuned for example configurations and tutorials covering how this all works. :)
Tuesday, May 19, 2009
More profiling results
So whilst I wait for some of the base restructuring in LUSCA_HEAD to "bake" (ie, settle down, be stable, etc) I've been doing some more profiling.
Top CPU users on my P3 testing box:
root@jennifer:/home/adrian/work/lusca/branches/LUSCA_HEAD/src# oplist ./squid
CPU: PIII, speed 634.464 MHz (estimated)
Counted CPU_CLK_UNHALTED events (clocks processor is not halted) with a unit mask of 0x00 (No unit mask) count 90000
samples % image name symbol name
1832194 8.7351 libc-2.3.6.so _int_malloc
969838 4.6238 libc-2.3.6.so memcpy
778086 3.7096 libc-2.3.6.so malloc_consolidate
647097 3.0851 libc-2.3.6.so vfprintf
479264 2.2849 libc-2.3.6.so _int_free
468865 2.2354 libc-2.3.6.so free
382189 1.8221 libc-2.3.6.so calloc
326540 1.5568 squid memPoolAlloc
277866 1.3247 libc-2.3.6.so re_search_internal
256727 1.2240 libc-2.3.6.so strncasecmp
249860 1.1912 squid httpHeaderIdByName
248918 1.1867 squid comm_select
238686 1.1380 libc-2.3.6.so strtok
215302 1.0265 squid statHistBin
Sigh. snprintf() leads to this:
root@jennifer:/home/adrian/work/lusca/branches/LUSCA_HEAD/src# opsymbol ./squid snprintf
CPU: PIII, speed 634.464 MHz (estimated)
Counted CPU_CLK_UNHALTED events (clocks processor is not halted) with a unit mask of 0x00 (No unit mask) count 90000
samples % image name symbol name
-------------------------------------------------------------------------------
15840 1.7869 squid safe_inet_addr
29606 3.3398 squid pconnPush
54021 6.0940 squid httpBuildRequestHeader
64916 7.3230 libc-2.3.6.so inet_ntoa
105692 11.9229 squid clientSendHeaders
127929 14.4314 squid urlCanonicalClean
196481 22.1646 squid pconnKey
265474 29.9476 squid urlCanonical
36879 100.000 libc-2.3.6.so snprintf
813284 91.7449 libc-2.3.6.so vsnprintf
36879 4.1602 libc-2.3.6.so snprintf [self]
12516 1.4119 libc-2.3.6.so vfprintf
9209 1.0388 libc-2.3.6.so _IO_no_init
Double sigh. Hi, the 90's called, they'd like their printf()-in-performance-critical code back.
Subscribe to:
Posts (Atom)





