04 April 2016

ZFS vs Hardware Raid System, Part III

As it was found in a previous post that the read/write rate varied a lot between the different controllers, the read/write tests need to be redone.
In the previous test, 10GB sized files have been used since that's more at the order of file size when used as GridPP storage behind DPM.  However, since the machine used for the tests has 24GB RAM which is larger than the file size, the files could still be in the cache.

For the following test, the same machine was used as in the above mentioned post, but with some changes in the configuration:

  1. Both raid controllers, a PERC H700 and a PERC H800, have been reset before doing any  tests.
  2. The element size defined in the controllers was 8KB for the previous test, now it uses 64KB.
  3. The test file size was increased to 30GB to be larger than the total RAM in the machine.
  4. All write and read tests were repeated 10 times to see how large the variation in the measured rates are.
  5. On the H700 11x2TB disks are used as raid6/raidz2 + 1 hotspare, mounted under /tank-2TB.
  6. On the H800 16x8TB disks are used as raid6/raidz2 + 1 hotspare, mounted under /tank-8TB.
The controller cache was again set to "write through" instead of the default "write back".
All read and writes have been performed 10 times to 10 different files to reduce the possibility that anything is left over in memory or controller cache. The results are then averaged for  read/write operations and controller. "dd" was used to generate/read the files with commands like:

time (dd if=/dev/zero of=/tank-2TB/test30G-$i bs=1M count=30720 && sync)
time (dd if=/tank-2TB/test30G-$i of=/dev/null bs=1M && sync)

The averaged results are given as "time"-value because the value reported by time includes the "sync" operation and makes sure everything is written to disk, while "dd" reports only about the own process which doesn't mean it is physically on disk already but can still be in memory cached. The minimum and maximum values are also reported to give an idea about the range during all 10 trials.

H700

ZFS write:                    56s (549MB/s) (min:  52s, max:  59s)
Hardware raid write:  305s (101MB/s) (min:265s, max:342s)

ZFS read:                   74s (415MB/s) (min:  56s, max:   83s)
Hardware raid read: 156s (197MB/s) (min:147s, max: 159s)

H800

ZFS write:                   28s (1097MB/s) (min:  28s, max:  30s)
Hardware raid write: 147s (  209MB/s) (min:125s, max:154s)

ZFS read:                     30s (1024MB/s) (min:30s, max: 34s)
Hardware raid read:   29s (1059MB/s) (min:29s, max: 31s)


In conclusion, it can be seen that the H800 performs in both configurations better than the H700 while ZFS clearly has the better performance than the hardware raid configuration. Therefore, all new installations at the Edinburgh side will use ZFS for the administration of the GridPP storage space. In the next blog post, I will show how to setup the zfs storage part for GridPP usage.

31 March 2016

Some thoughts on data in academic environments vs industry - Part 2/2

Now if we in the academic world are so good at managing "big data," how does it (if at all) impact industry and society as a whole? Obviously we rely on the storage and networking industry for the "fabric," and more generally we widely collaborate with industry partners in "big data" projects such as those funded by Horizon 2020, e.g. ESiWACE and SAGE to solve the "next generation" of problems. And of course there are data specialists in industry with even bigger data volumes than ours.

So the question here is how we (=academic data centres) can engage with industry more widely, specifically with those who could benefit from the expertise we have developed.

If we have something that can be commercialised, we can of course spin out companies: universities and research councils can do this fairly easily. Indeed, many former collaborators are now CxOs and co-founders of startups such as SIXSQ and StreamVibe. Also patents and knowledge exchange. And STFC has the Hartree centre which focuses on solving problems - including "big data" ones - for industry. STFC also has an innovation hub.

For more collaborative exploratory work, let me highlight the opportunity of working with the Connected Digital Economy Catapult and the Big Innovation Centre. Like we see with our own experiments, we need to get experts together from both sides - experiment and infrastructure - to make them both work together. CDEC and BIC have the ability to tap into business requirements, and as an example, we have previously investigated and designed a "trusted data hub" to enable companies that don't want to share data openly to share in a controlled way within a trusted platform. By connecting academic expertise and innovation with industry requirements, we can increase the impact of our work and together design systems which improve people's lives. The CDEC and BIC wear the suits, so we don't have to!

29 March 2016

Deletion of Tape backed data for ATLAS VO at RAL Tier1

The RAL Tier1 is just about to migrate all the tape backed data that we store for the ATLAS collaboration onto larger tapes ( well the tapes are the same size but the amount per tape which can be written with new drives is higher. Before we started; we asked ATLAS if they could find any data to delete before we migrate ( so a not to put gaps  in to the new tapes.)
This they did successfully.
They deleted 1.58 million files; (corresponding to 1.48 PB of data,) over a 5 day  period this can be seen in the following plots, (deletion rate when busy of 20k files per hour):






Now to just get ATLAS to delete the remaining logfiles from DATADISK... (requires moving  to using our new CEPH storage system for this to happen...)....

24 March 2016

Some thoughts on data in academic environments vs industry, part 1 (of 2)

I was asked today about my opinion on the difference between big data in academia and industry. As I see it we have volume (data collections on the order of 10s of PBs, e.g. WLCG, climate); characteristically velocity is an order of magnitude greater than volume as all data is moved around and replicated (FTS alone was recently estimated to have moved nearly a quarter of an exabyte in a year, and Globus advertise their (presumably estimated) data transfer volumes on their home page) Most science data is copied and replicated, as large scale data science is a global endeavour, requiring the collaboration of research centres across the world.

But we (science) have less variety. Physics events are physics events, and even with different types like raw, AOD, ESD, etc., there is a manageable collection of formats. Different communities have different formats, but as a rule, science is fairly consistent.

Bandwidth into sites is measured in 10s of Gb/s (10, 40, 60, that sort of stuff); and the 16 PB Panasas system for JASMIN can shift something like 2.5 Tb/s (think of it as a million times faster than your home Internet)

Moreover, for some things like WLCG, data models are very regimented, thus ensuring that the processing happens in an orderly fashion rather than chaotic. We have security strong enough to enable us to run services outside the firewall (as otherwise we'd slow the transfer down a lot, and/or melt the firewall).

And the expected evolution is more of the same - towards the 100PB for data collections, 1EB for data transfers. More data, more bandwidth, perhaps also more diversity within research disciplines. When will we get there? If growth is truly exponential, it could be too soon  :-) we'd get more data before the technology is ready... even with sub-exponential growth, we may have scalability issues - sure, we could store an exabyte today in a finite footprint, but can we afford it?

"Big Science" security goals are different - data is rarely personal, but may be commercially sensitive (e.g. MX of a protein) or embargoed for research reasons. Integrity is king, availability is the prince, and the third "classic" data security goal, confidentiality, not quite a pauper, but a baronet, maybe?! It's an oversimplified view but these are often the priorities. Big Science requires public funding so the data that comes out of it must be shared to maximise the benefit.

There's also the human side. With the AoD (Archer-on-Duty) looking after our storage systems, it does not seem likely we will ever have ten times more people for an order of magnitude more data.  But so far, comparing to when data was order-of-1PB scale, we didn't need ten times more people than we had then.

Agree? Disagree? Did I forget something?

23 March 2016

ZFS vs Hardware Raid System, Part II

This post will focus on other differences between a ZFS based software raid and a hardware raid system that could be important for the usage as GridPP storage backend. In a later post, the differences in the read/write rates will be tested more intensively.
For the current tests, the system configuration is the same as described previously

First test is, what happens if we just take a disk out of the current raid system...
In both cases, the raid gets rebuild using the hot spare that was provided in the initial configuration. However, the times needed to fully restore redundancy are very different:

ZFS based raid recovery time: 3min
Hardware based raid recovery time: 9h:2min

For both systems, only the test files from the previous read/write tests were on disk, and the hardware raid was initialized newly to remove the corrupted filesystem after the failure test and then the test files were recreated.  Both systems where not doing anything else during the recovery period.

The large difference is due to the fact that ZFS is raid system, volume manager, and file system all in one. ZFS knows about the structure on disk and the data that was on the broken disk. Therefore it only needs to restore what actually was really used, in the above test case just only 1.78GB.
The hardware raid system on the other hand knows nothing about the filesystem and real used space on a disk, therefore it needs to restore the whole disk even if it's like in our test case nearly empty - that's 2TB to restore while only about 2GB are actually useful data! This difference will  become even more important in the future when the capacity of a single disk gets larger.

ZFS is now also used on one of our real production systems behind DPM. The zpool on that machine, consist of 3 x raidz2 with 11 x 2TB disks for each raidz2 vdev, also had a real failed disk which needed to be replaced. There are about 5TB of data on that zpool, and the whole time needed to restore full redundancy took about 3h while the machine was still in production and used by jobs to request or store data.  This is again much faster than what a hardware raid based system would need even if the machine would be doing nothing else in the meantime than restoring full redundancy.



Another large difference between both system is the time needed to setup a raid configuration. 
For the hardware based raid, first one need to create the configuration in the controller, initialize the raid, then create partitions on the virtual disk, and lastly format the partitions to put a file system on it. When all is finished, the newly formatted partitions need to be put into /etc/fstab to be mounted when the system starts, mount points need to be created, and then the partitions can be mounted. That's a very long process which takes up a lot of time before the system can be used.
To give an idea about it, the formatting of a single 8TB partition with ext4 took
on H700: about 1h:11min
on H800: about       34min

In our configuration, 24x2TB on H800 and 12x2TB on H700, the formatting alone would take about 6h! (if done one partition after another)

For ZFS, we still need to create a raid0 for each disk separately in the H700 and H800 since both controllers don't support JBODs. However once this is done, it's very easy and fast to create a production ready raid system. There is one single command which does everything: 
zpool create NAME /dev/... /dev/... ..... spare /dev/....
After that single command, the raid system is created, formatted to be used, and mounted under /NAME which can also be changed using options when the zpool is created (or later). There is no need to edit /etc/fstab and the whole setup takes less than 10s
To have single 8TB "partitions", one can create other zfs in this pool and set a quota for it, like
zfs -o refquota=8TB create NAME/partition1
After that command, a new ZFS is created, a quota is placed on it which makes sure that tools like "df" only see 8TB available for usage, and it's mounted under /NAME/partition1 - again no need to edit /etc/fstab and the setup takes just a second or two. 




Another important consideration is what happens with the data if parts of the system fail. 
In our case, we have 12 internal disks on a H700 controller and 2 MD1200 devices with 12 disks each that are connected through a H800 controller, and 17 of the 2TB disks we used so far in the MD devices need to be replaced by 8TB disks.  There are different setups possible and it's interesting to see what in each case happens if a controller or one of the MD devices fails.
The below mentioned scenarios assume that the system is used as a DPM server in production.

Possibility 1: 1 raid6/raidz2 for all 8TB disks, 1 raid6/raidz2 for all 2TB disks on the H800, and 1 raid6/raidz2 for all 2TB disks on the H700

What happens if the H700 controller fails and can't be replaced (soon)?

Hardware raid system
The 2 raid systems on H800 would still be available. The associated filesystems could be drained in DPM, and the disks on the H700 controller could be swapped with the disks in the MD devices associated with the drained file systems. However, if the data on the disks will be usable depends a lot on the compatibility of the on disk format used by different raid controllers. Also, even if that would be possible to recreate the raid on different controller without any data lost (which is probably not possible), it would appear to the OS as a different disk and all the mount points would need to be changed to have the data available again.  

ZFS based raid system
Again, the 2 raid systems associated with the H800 would still be available and the pool of 8TB disks could be drained. Then the disks on the failed H700 could simply be swapped with the 8TB disks, and the zpool would be available again - mounted under the same directories before no matter in which bays the disks are or on which controller.

What happens if one of the MD devices fails?

Hardware raid system
Since one MD device has only 12 disks and we need to put in 17 x 8TB disks, one raid system consist of disks from both MD devices. If any MD device fails, the data on all 8TB disks will be lost. If the MD device with the remaining 2 TB disks fails, it depends if the raid on the 2TB disks could be recognized by the H700 controller. If so, then the data on the disks associated with the H700 could be drained and at least the data from the 2TB disks on the failed MD device could be restored, although with some manual configuration changes in the OS since DPM needs the data mounted in the previously used directories.

If anyone has experience on the possibility to use disks from a raid created on one controller  on another controller of a different kind, I would appreciate if we could read about it in the comments section.


ZFS based raid system
If any MD device fails, the pool with the internal disks on the H700 could be drained, and then the disks swapped with the disks on the failed MD device. All data would be available again and no data would be lost.

One raidz2 for all 8TB disks (pool1) and one raidz2 for all 2TB disks (pool2)

This setup is only possible when using a ZFS based raid since it's not supported by hardware raid controllers.

If the H700 fails, then the data on pool1 could be drained and the 8TB disks be replaced with the 2 TB disks which were connected to the H700. Then all data would be available again.
It's similar when one of the MD devices fails. If the MD device with only 8TB disks fails, then pool2 can be drained and the disks on the H700 be replaced with the 8TB disks from the failed MD device. 
If the MD device with 2TB and 8TB disks fails, it's a bit more complicated but all data could be restored in a 2 step process: 
1) replace the 2TB disks on the H700 with the 8TB disks from the failed MD device which makes pool1 available again which then could be drained.
2) Put the original 2TB disks that were connected to the H700 back in and replace 8TB disks on the working MD device with the remaining 2TB disks from the failed MD device, which makes pool2 available again.
In any case, no data would be lost.




22 March 2016

ZFS vs Hardware Raid

Due to the need of upgrading our storage space and the fact that we have in our machines 2 raid controllers, one for the internal disks and one for the external disks, the possibility to use a software raid instead of a traditional hardware based raid was tested.
Since ZFS is the most advanced system in that respect, ZFS on Linux was tested for that purpose and proved to be a good choice here too.
This post will describe the general read/write and failure tests, and a later post will describe additional tests like rebuilding of the raid if a disk fails, different failure scenarios, setup and format times.
Please, use the comment section if you would like to have other tests done too.


harware test configuration:

  1. DELL PowerEdge R510 
  2. 12x2TB SAS (6Gbps) internal storage on a PERC H700 controller
  3. 2 external MD1200 devices with 12x2TB SAS (6Gbps)on a PERC H800 controller
  4. 24GB RAM
  5. 2 x Intel Xeon E5620 (2.4GHz)
  6. for all settings in the raid controllers the default was used for all tests, except for cache which was set to "write through"
ZFS test system configuration:
  1. SL6 OS
  2. ZFS based on the latest version available in the repository 
  3. no ZFS compression used
  4. 1xraidz2 + hotspare for all the disks on H700  (zpool tank)
  5. 1xraidz2 + hotspare for all the disks on H800  (zpool tank800)
  6. in both raid controllers each disk is defined as a single raid0 since they don't support JBOD, unfortunately
Hardware raid test system configuration:
  1. same machine with same disks, controllers, and OS used as for the ZFS test configuration
  2. 1xraid6 + hotspare for all the disks on H700
  3. 1xraid6 + hotspare for all the disks on H800 
  4. space was divided into 8TB partitions and formatted with ext4

Read/Write speed test



  • time (dd if=/dev/zero of=/tank800/test10G bs=1M count=10240 && sync)
  • time (dd if=/tank800/test10G of=/dev/null bs=1M && sync)
  • first number in the results is given by "dd"
  • time and second number is given by "time"
  • write test was done first for both controllers, and then the read tests


  • H700 results

    ZFS based:
    write: 236MB/s, 1min:02 (165MB/s)
    read:  399MB/s, 0min:27 (379MB/s)

    Hardware raid based:
    write: 233MB/s, 1min:10 (146MB/s)
    read:    1.2GB/s, 0min:18 (1138MB/s)

    H800 results

    ZFS based:
    write: 619MB/s, 0min:23 (445MB/s)
    read:  2.0GB/s, 0min:05 (2048MB/s)

    Hardware raid based:
    write: 223MB/s, 1min:13 (140MB/s)
    read:  150MB/s, 1min:12 (142MB/s)

    H700 and H800 mixed

    • 6 disks from each controller were used together in a combined raid configuration
    • this kind of configuration is not possible for a hardware based raid
    ZFS result:
    write: 723MB/s, 0min:37 (277MB/s)
    read:  577MB/s, 0min:18 (568MB/s)

    Conclusion

    • ZFS rates for H800 based raid much better than hardware raid based system
    • the large difference between  ZFS and hardware raid based reads needs more investigation
      • for repeating the same tests 2 more times it was at the same order, however
    • H800 has a much better performance than H700 when using ZFS, but not for the hardware raid configuration

    Failure Test

    Here it was tested what happens if a 100GB file (test.tar) is copied (cp and rsync) from the H800 based raid to the H700 based raid and during this copy the system failed, simulated by cold reboot through the remote console.

    ZFS result:

    root@pool6 ~]# ls -lah /tank 
    total 46G
    drwxr-xr-x.  2 root root    5 Mar 19 20:11 .
    dr-xr-xr-x. 26 root root 4.0K Mar 19 20:17 ..
    -rw-r--r--.  1 root root  16G Mar 19 19:07 test10G
    -rw-r--r--.  1 root root  13G Mar 19 20:12 test.tar
    -rw-------.  1 root root  18G Mar 19 20:06 .test.tar.EM379W

    [root@pool6 ~]# df -h /tank
    Filesystem      Size  Used Avail Use% Mounted on
    tank             16T   46G   16T   1% /tank
    [root@pool6 ~]# du -sch /tank
    46G     /tank
    46G     total

    [root@pool6 ~]# rm /tank/*test.tar*
    rm: remove regular file `/tank/test.tar'? y
    rm: remove regular file `/tank/.test.tar.EM379W'? y
    [root@pool6 ~]# du -sch /tank
    17G     /tank
    17G     total

    [root@pool6 ~]# ls -la /tank
    total 16778239
    drwxr-xr-x.  2 root root           3 Mar 19 20:21 .
    dr-xr-xr-x. 26 root root        4096 Mar 19 20:17 ..
    -rw-r--r--.  1 root root 17179869184 Mar 19 19:07 test10G
    • everything consistent
    • no file check needed at reboot
    • no problems at all occurred 

    Hardware raid based result:

    [root@pool7 gridstorage02]# ls -lhrt
    total 1.9G
    drwx------    2 root   root    16K Jun 26  2012 lost+found
    drwxrwx---   91 dpmmgr dpmmgr 4.0K Feb  4  2013 ildg
    -rw-r--r--    1 root   root      0 Mar  6  2013 thisisgridstor2
    drwxrwx---   98 dpmmgr dpmmgr 4.0K Aug  8  2013 lhcb
    drwxrwx---  609 dpmmgr dpmmgr  20K Aug 27  2014 cms
    drwxrwx---    6 dpmmgr dpmmgr 4.0K Nov 23  2014 ops
    drwxrwx---    6 dpmmgr dpmmgr 4.0K Mar 13 12:18 ilc
    drwxrwx---    9 dpmmgr dpmmgr 4.0K Mar 13 23:04 lsst
    drwxrwx---  138 dpmmgr dpmmgr 4.0K Mar 14 10:23 dteam
    drwxrwx--- 1288 dpmmgr dpmmgr  36K Mar 15 00:00 atlas
    -rw-r--r--    1 root   root   1.9G Mar 18 17:11 test.tar

    [root@pool7 gridstorage02]# df -h .
    Filesystem            Size  Used Avail Use% Mounted on
    /dev/sdb2             8.1T  214M  8.1T   1% /mnt/gridstorage02

    [root@pool7 gridstorage02]# du . -sch
    1.9G    .
    1.9G    total

    [root@pool7 gridstorage02]# rm test.tar 
    rm: remove regular file `test.tar'? y

    [root@pool7 gridstorage02]# du . -sch
    41M     .
    41M     total

    [root@pool7 gridstorage02]# df -h .
    Filesystem            Size  Used Avail Use% Mounted on
    /dev/sdb2             8.1T -1.7G  8.1T   0% /mnt/gridstorage02

    • Hardware raid based tests were done first, on a  machine that was previously used as dpm client, therefore the directory structure was left, but empty
    • during the reboot a file system check was done
    • "df"  reports a different number for the used space than "du" and "ls"
    • after removing the file, the used space reported by "df" is negative
    • file system is not consistent anymore

    Conclusion here:

    • for the planned extension (17x2TB exchanged for 8TB disks), the new disks should be placed in the MD devices and managed by the H800 using ZFS
    • second zpool can be used for all remaining 2TB disks (on H700 and H800 together)
    • ZFS seems to handle system failures better 
    To be continued...

    27 January 2016

    Xrootd for all

    Xrootd provides, amongst other things, a convenient method to externally access files at a site anywhere in the world using your grid credentials.

    Specialist storage systems such as DPM and dcache now include a xrootd server in their deployment.
    If you use a stranded POSIX file system (e.g. Lustre, GPFS, NFS) it's possible to set up a standalone xrootd server to export all or part of the file system to external clients.

    The LHC experiments have have gone a step further and have set up federated storage services combining storage from several separate sites in to one namespace allowing seamless client access to storage without having to worry where the data is stored. They have provided instructions to setup a service but only for their VO e.g.



    Extending this to allow access to other VOs data via xrootd but without the federated storage service is simple.

    The xrootd server runs as user xrootd. In order to access files it must have the correct permissions to the files. This can be done by making the xrootd user a member of the appropriate groups across the site (e.g. via NIS).

    ypcat -k group 
    ...
    dteam dteam:x:12345:user1,user2, …, xrootd
    atlas atlas:x:13345:user1,user2, …, xrootd

    For simplicity I'll make a symlink to the file system I want to export on the xrootd server, e.g.

    ln -sf /mnt/lustre_2/storm_3/atlas/ /atlas
    ln -sf /mnt/lustre_2/storm_3/dteam/ /dteam

    The xrootd server configuration file is /etc/xrootd/xrootd-clustered.cfg. Within this file we need to define the file system to export and do so read only for security,

    all.export /dteam r/o
    all.export /atlas r/o

    We also need to add the VOs to the X509 configuration.

    ...
    sec.protparm gsi -vomsfun:/usr/lib64/libXrdSecgsiVOMS.so -vomsfunparms:certfmt=raw | vos=atlas,dteam | grps=/atlas,/dteam
    acc.authdb /etc/xrootd/auth_file
    ...

    The /etc/xrootd/auth_file specifies the group/user access rights. The following will give read and list right to members of the atlas group to file under /atlas and dteam group members for files under /dteam

    g /atlas /atlas rl
    g /dteam /dteam rl


    The final configuration files look like

    cat xrootd-clustered.cfg
    ...
    frm.xfr.copycmd /bin/cp /dev/null $PFN

    # atlas redirection
    all.manager atlas-xrd-uk.cern.ch+:1098
    xrootd.redirect atlas-xrd-uk.cern.ch:1094 ? /atlas
    all.sitename SITENAME

    all.export /dteam r/o
    all.export /atlas r/o

    all.role server
    all.adminpath /var/run/xrootd
    all.pidpath /var/run/xrootd
    xrootd.async off

    # atlas Monitoring
    if exec xrootd
    xrd.report atl-prod05.slac.stanford.edu:9931 every 60s all -buff -poll sync
    fi
    # if your sites is in EU uncomment the next line
    xrootd.monitor all flush 30s window 5s fstat 60 lfn ops xfr 5 dest redir files info user atlas-fax-eu-collector.cern.ch:9330

    # N2N configuration. Please change for your site
    oss.namelib /usr/lib64/XrdOucName2NameLFC.so

    # X509 configuration, change nothing
    xrootd.seclib /usr/lib64/libXrdSec.so
    sec.protparm gsi -vomsfun:/usr/lib64/libXrdSecgsiVOMS.so -vomsfunparms:certfmt=raw|vos=atlas,dteam|grps=/atlas,/dteam
    sec.protocol /usr/lib64 gsi -ca:1 -crl:3 -gridmap:/dev/null
    acc.authdb /etc/xrootd/auth_file
    acc.authrefresh 60
    ofs.authorize

    [root@xrootd02 xrootd]# cat auth_file
    g /atlas /atlas rl
    g /dteam /dteam rl


    Note:
    Additional file systems can be added in the same fashion for more VOs e.g. snoplus, t2k …..

    It's possible to use argus server for the authentication http://londongrid.blogspot.co.uk/2014/10/xrootd-and-argus-authentication.html

    The bandwidth to the file system will be limited by the performance of the xrootd server. For local file access it's still better to use native POSIX access, especially with parallel file systems like Lustre.



    04 January 2016

    Update on vo.dirac.ac.uk data movement and filesize distribution.

    So....... I should have known that the information I posted in the blog post in November of last year would soon be out of date; but I didn't think it would be this soon! DiRAC have successfully developed their system to tar and split their data samples before transferring into the RAL Tier1. This system has dramatically increased the data transfer rates.
     What  has also changed is the number of files per tape  due to the change in average filesize per tape:
     
    This has meant the number of files per tape varied from a starting value of 2-3 thousand per tape , swelling top 2-3 million before finally settling on 20-40 per tape. ( file size is ~ 250-300GB per file.)
    To move large files requires good transfer rates; which we have been able to achieve; (can be seen in this log snippet):

    Tue Dec 29 08:00:28 2015 INFO     bytes: 293193121792, avg KB/sec:286321, inst KB/sec:308224, elapsed:1001
    Tue Dec 29 08:00:33 2015 INFO     bytes: 294824181760, avg KB/sec:286481, inst KB/sec:318566, elapsed:1006
    Tue Dec 29 08:00:38 2015 INFO     bytes: 296458387456, avg KB/sec:286643, inst KB/sec:319180, elapsed:1011
    Tue Dec 29 08:00:43 2015 INFO     bytes: 298053795840, avg KB/sec:286766, inst KB/sec:311603, elapsed:1016
    Tue Dec 29 08:00:45 2015 INFO     bytes: 298822410240, avg KB/sec:286715, inst KB/sec:268071, elapsed:1018


    Incidentally, the large filesize also helps reduce the overall rate loss due to individual overhead setup and completion per transfer. ( overhead of ~15 seconds for this file which then took 1018 seconds to transfer.This has allowed us to transfer ~ 125Tb of data over the new year period:

    And a completion rate of ~90%

    Although the low number of transfers does not allow the FTS optimizer to change settings so as to improve the throughput rate:


    Let's hope we can continue this rate. My next step is to look at the rate at which we can create the tarballs on the source host in preparation for transfer  and whether this technique can be applied at other source sites within vo.dirac.ac.uk.

    28 December 2015

    There is no such thing as a "Solution"

    As we near the end of 2015 and look forward to 2016, it is time to reflect on the past year and look ahead to the next. The GridPP infrastructure can clearly deal with the tens/hundreds petabyte range but the middleware is just that, "middle." Some work always is needed to adapt to every new community's requirements. The large LHC experiments can move away from the established codebase and specialise their infrastructure which is both a good thing (infrastructure is flexible) and bad (not all communities can roll their own infrastructure). Using open and interoperable standards helps.

    There exists research which is not high energy physics (HEP). These people look at the HEP infrastructure - the global infrastructure that found the Higgs - and say "we don't work like that."

    Yet there is a role for GridPP's storageanddatamanagement infrastructure and expertise - most research areas of have data, and growing volumes of it. The expertise in moving, storing, and making data discoverable and available for processing is important regardless of which area of research the data belongs to. It doesn't matter that their analysis is different; their data problems are the same.

    Big data volumes are also already found in astronomy, bioinformatics, Earth sciences, etc., and these communities have developed methods for data management, too. So 2016 will be an important year to continue the discussions with their infrastructure managers and developers - while every community is different and there is no such thing as a "solution," the more expertise and infrastructure we can share, the more easily we can manage the research data challenges of the future.

    16 December 2015

    DPM Workshop 2015

    This year, the annual DPM Collaboration Workshop was hosted by the core DPM development team themselves, at CERN.

    As usual, it was a two-day affair, headed by regional site reports from around the world - Italy, France, Australia and the UK reported.

    One common topic between the site reports was discussion of the different approaches to configuring/maintaining DPM instances - France being a Quattor-dominant zone, with separately maintained configuration, Italy using a Foreman-driven puppet (requiring some modification of the "standard" DPM puppet scripts), and Australia and the UK using the standard puppet local config or hand-configured systems.
    Another common topic was discussion of the recent ATLAS-driven storage tasks - the decommissioning of the large amount of data in PRODDISK, and the push for the new data dumps for consistency checking. People were generally rather unhappy about the PRODDISK task, especially in how slow it was to actually delete all the data from the token using the rfio-based admin tools. The newer davix-tools, which should be much faster, are unfortunately much less well known at present.

    (Several sites also called out Andrey Kirianov's dpm-dbck tool for deep database linting and consistency checking as being especially useful - and he had a talk later in the workshop describing his future release plans, incorporating feedback.)


    The Core DPM updates followed, beginning with the news that the DPM (and thus also LFC) team is being brought back into the fold of CERN IT-DSS, where they'll join the EOS, Castor and FTS teams. The DPM devs were each at pains to indicate that this would have no visible effects in the "short term".
    As for the future plans of the group, they focussed much on the well-known existing topics of the last year: removing the dependance on rfio, and srm, and leveraging this freedom to improve other aspects of the system (for example, better support for multiple checksums on files).
    Already, DPM 1.8.10 provides the "beta" versions of the new DPM Space Reporting functionality (which is intended to supplant the use of SRM for space reporting), although it is currently off by default (and is unable to track storage changes made via SRM). This functionality also broke the beta EGI Storage Accounting tools, until I submitted a patch to stop directory sizes being counted as well as the size of the individual files contained (essentially n-ple counting every file in the final tally, with n the directory depth of the file!).
    GridFTP Redirection is another feature which the DPM team were keen to point to as a functioning enhancement in DPM 1.8.10. As with the Space Reporting, however, it comes with caveats: the enhanced performance gained from handing off from the Head nodes' GridFTP service to the correct pool nodes' GridFTP as part of the handshake is only possible for clients which support the Delayed PASV operation mode. Neither gfal1 (ie lcg-cp) nor uberftp support this, causing the new GridFTP to fall back to a much less efficient mechanism, which actually has worse network characteristics than the old GridFTP. Additionally, as this mechanism depends on patching GridFTP itself, new releases are needed every time a new Globus release happens...

    The focus of the DPM core in the next year is to be removing the rfio underpinnings of DPM, which are used for all of the low-level communication between the head and pool nodes, in favour of a new, RESTful interface based on FastCGI. This is based on some exploratory work by Eric Cheung, and is also planned to bring in additional enhancements for non-SRM DPMs (for example, fully supported space management with directory quotas).

    While I did the UK report, Bristol's Luke Kreszko headed up the afternoon's talks with a great talk on the experience of running a lightweight "dmlite" style DPM on HDFS, without an SRM. While Luke had some bugs to report, he was keen to make clear that Andrea and the rest of the DPM core team had responded extremely quickly to each bug report and snag, and that the system was improving over time. (Outstanding issues include a lack of a way to throttle load on any particular service (which is also a wider concern of the UK), and a desire to have more intelligent selection of servers for each request.)

    There was also an interesting series of talks on the HTTP provision for Experiments - firstly, from the perspective of the HTTP Deployment Task Force, which exists to manage and enable the deployment of HTTP/WebDAV interfaces at sites with the functionality needed to support the Experiments, and secondly from the announcement of an initiative to attempt to run ATLAS work against a "pure" HTTP DPM. (This was a little short of the promise of the title, as data transfers were envisaged to occur over GridFTP, using the new GridFTP redirection, rather than HTTP.)

    We also learned, from the Belle 2 Experiment, that they were currently rolling out their storage and data management infrastructure based entirely around DPM and the LFC as a file catalog!

    Heading up the day, there was also a talk on the CMS "performance testing" for DPM sites using the CMS AAA Xrootd Federation. I have to say, I wasn't entirely convinced by some of the "improvements" demonstrated by the graphs, as the "new" tests seemed to stop before approaching the kinds of load which were problematic for the "old" tests. However, it was an interesting insight into CMS's thinking about dealing with the fact that not all sites will meet their performance standards for AAA, with the idea of two Federations (a "proper" AAA, and a "not as good" AAA) existing, with sites being shunted transparently from the "proper" one to the backup if their performance characteristics dropped sufficiently to compromise the federation.

    The rest of the meeting was mostly feedback and live demos of the Puppet configuration and existing tools, which was useful and involving.

    23 November 2015

    First Analysis of file size distribution for vo.dirac.ac.uk at RAL Tier1.

    SO the DiRAC virtual organization has successfully  moved ~4M files (95TB ) of data to the RAL Tier1. But in what form:
    Smallest file size is ( ignoring the 40k zero size files) is 2B
    Median file size is 450kB
    The modal average (114k files out of 4M)  is 1.4MB
    Mean file size is 23.15MB
    Largest file size is 482GB

    Looking at the file size distribution is interesting.

     The VO hope to improve this by "tar"ing and compressing the files; (ideally we would like the files to be ~1Gb in size.)

    14 October 2015

    Dave's locations in Rucio

    ATLAS have now fully moved over to their new data management system RUCIO and have also been consolidating  the locations of me and my offspring. Currently the situation is as follows:


    I have have  697 unique children of which:
    541 have no Clones.
    129 have 1 Clone.
    27 have 2 Clones ( such as me).
    And 1 has 3 clones.
    Compare this to early on in my life when some datasets had over 25 clones!!!

    There are 637 "rooms" spread over 136 "houses" in total. I and my progeny are in 40 houses spread over 70 rooms. The types of room are:
         22 LOCALGROUPDISK
         16 DATADISK
          6 PERF-MUONS
          5 PHYS-SM
          5 PERF-EGAMMA
          5 DATATAPE
          4 PHYS-SUSY
          2 PHYS-BEAUTY
          1 TZERO
          1 SCRATCHDISK
          1 PERF-JETS
          1 PERF-FLAVTAG
          1 DET-LARG

    The type of children I have breaks down as following:
    607 Dirk's, 13 Gavin's and  79 Ursula's.

    This houses are very international being spread over.
    Canada
    Czech Republic
    France
    Germany
    Italy
    Japan
    Netherlands
    Nordic Countries
    Portugal
    Russia
    Spain
    Switzerland
    Taiwan
    Turkey
    UK
    USA

    25 September 2015

    Milestone Passed with Last Two Years of Transfers Through FTS

    Over the Last Two years; the FTS system (as used by the WLCG VOs and others) have moved ~0.5EB of data (over a Billion files). What is an EB ? Well its 1000 PBytes or 10^6 TBytes or 10^9 GBytes or 10^12MBytes or 10^15kBytes or 10^18Bytes. This can be seen form the monitoring page: