03 February 2014

Its been a while, but my family dynamics are changing....

Due to housing and room capacity; ATLAS decided to reduce the number of centrally controlled clones. So here is an update for where Georgina, Eve and I are now living after ATLAS's change of policy:

New Values
DataSet Name
Dave G'gina Eve
"DNA" Number
Number of "Houses" 49 70 54
Type of Rooms:DATADISK
17 37 33
Type of Rooms:LGD
32 58 24
Type of Rooms:PERF+PHYS
29 56 24
Type of Rooms:TAPE
7 12 9
Type of Rooms:USERDISK
0 1 5
Type of Rooms:CERN
8 10 10
Type of Rooms:SCRATCH
3 6 1
Type of Rooms:CALIB 0 4 7
Total number of people (including clones) 1090 1392 471
Number of unique people 876 1019 293
Numer of "people" of type:
^user 136 368 97
Numer of unique "people" of type:
^user 131 340 83
Numer of "people" of type:
^data 919 950 352
Numer of unique "people" of type:
^data 719 616 189
Numer of "people" of type:
^group 34 74 22
Numer of unique "people" of type:
^group 25 63 21
Numer of "people" of type:
^valid 1 0 0
Numer of unique "people" of type:
^valid 1 0 0
 Datasets that have 1 copy 696 763 184
 Datasets that have 2 copies 146 184 62
 Datasets that have 3 copies 34 44 33
 Datasets that have 4 copies 0 16 9
 Datasets that have 5 copies 0 7 4
 Datasets that have 6 copies 0 5 0
 Datasets that have 7 copies 0 0 0
 Datasets that have 8 copies 0 0 1
 Datasets that have 12 copies 0 0 0
 Datasets that have 13 copies 0 0 0
Number of files that have  1 copy 56029 134672 14260
Number of files that have  2 copies 8602 35191 6582
Number of files that have  3 copies 1502 1751 5924
Number of files that have  4 copies 0 1879 75
Number of files that have  5 copies 0 607 868
Number of files that have  6 copies 0 306 0
Number of files that have  7 copies 0 0 0
Number of files that have  8 copies 0 0 1
Number of files that have  12 copies 0 0 0
Number of files that have  13 copies 0 0 0
Total number of files on the grid: 77739 222694 49844
Total number of unique files: 66133 174406 27710
Data Volume (TB) that has  1 copy 9.175 30.389 4.127
Data Volume (TB) that has  2 copies 6.361 21.62 3.294
Data Volume (TB) that has  3 copies 0.223 2.452 7.812
Data Volume (TB) that has  4 copies 0 1.263 0.14
Data Volume (TB) that has  5 copies 0 1.408 0.17
Data Volume (TB) that has  6 copies 0 0.379 0
Data Volume (TB) that has  7 copies 0 0 0
Data Volume (TB) that has  8 copies 0 0 0.001
Data Volume (TB) that has  12 copies 0 0 0
Data Volume (TB) that has  13 copies 0 0 0
Total Volume of data on the grid (TB): 22.57 95.351 35.57
Total Volume of unique data (TB): 15.76 57.511 15.54



The difference in values from my last update are:


Difference
DataSet Name
D' G' E'
"DNA" Number


Number of "Houses" -10 -9 -9
Type of Rooms:DATADISK
-11 -12 -17
Type of Rooms:LGD
-5 -3 10
Type of Rooms:PERF+PHYS
-3 -2 2
Type of Rooms:TAPE
1 0 0
Type of Rooms:USERDISK
-1 -8 0
Type of Rooms:CERN
5 4 5
Type of Rooms:SCRATCH
3 -6 -13
Type of Rooms:CALIB 0 -1 0
Total number of people (including clones) -76 -202 -171
Number of unique people -18 -101 -6
Numer of "people" of type:
^user -1 -102 33
Numer of unique "people" of type:
^user -1 -89 28
Numer of "people" of type:
^data 194 -98 -186
Numer of unique "people" of type:
^data 187 -15 -23
Numer of "people" of type:
^group 3 4 -18
Numer of unique "people" of type:
^group -1 9 -11
Numer of "people" of type:
^valid 0 0 0
Numer of unique "people" of type:
^valid 0 0 0
 Datasets that have 1 copy 5 -48 6
 Datasets that have 2 copies 3 -13 -1
 Datasets that have 3 copies -18 -24 8
 Datasets that have 4 copies -71 -7 2
 Datasets that have 5 copies -1 0 -17
 Datasets that have 6 copies 0 0 -2
 Datasets that have 7 copies 0 -2 -1
 Datasets that have 8 copies 0 -1 0
 Datasets that have 12 copies 0 0 -6
 Datasets that have 13 copies 0 0 -3
Number of files that have  1 copy 2385 -6991 4751
Number of files that have  2 copies -4302 -1672 -1123
Number of files that have  3 copies -358 -1957 1404
Number of files that have  4 copies -7 -213 -35
Number of files that have  5 copies -7 35 -223
Number of files that have  6 copies 0 -181 -20
Number of files that have  7 copies 0 -142 -5
Number of files that have  8 copies 0 -73 0
Number of files that have  12 copies 0 0 -6
Number of files that have  13 copies 0 0 -5
Total number of files on the grid: -7356 -19547 26871
Total number of unique files: -2289 -11194 -16964
Data Volume (TB) that has  1 copy 0.375 2.689 2.427
Data Volume (TB) that has  2 copies -0.439 0.32 -1.406
Data Volume (TB) that has  3 copies -0.077 -3.148 0.612
Data Volume (TB) that has  4 copies -0.011 -0.237 0.02
Data Volume (TB) that has  5 copies -0.028 0.008 -0.3
Data Volume (TB) that has  6 copies 0 0.019 -0.039
Data Volume (TB) that has  7 copies 0 -0.13 -0.001
Data Volume (TB) that has  8 copies 0 -0.12 0
Data Volume (TB) that has  12 copies 0 0 -0.001
Data Volume (TB) that has  13 copies 0 0 -0.001
Total Volume of data on the grid (TB): -1.134 -8.649 -0.231
Total Volume of unique data (TB): -0.241 -0.489 1.344


The number of unique children in all three families (44TB/205k files) are completely unique to the world and are at risk to a single disk failure in a room. 

23 December 2013

Good talks at DPM workshop

It's nice to see plans and  issues for DPM sites are similar around the world. Here are my musings:

The fact Australia and Taiwan sites see the approximate percentage of dark files after ATLAS RUCIO renamed their site as I have seen in the UK (8-9%) is encouraging to know that the UK is not especially bad; (now just need to work with ATLAS on how to efficiently clean up these files, delete empty directories; and how to reduce dark data creation in the future.) It will be interesting to see if other VO's have a lower or higher percentage of dark data.

Also good to hear that a puppet deployment of DPM (rather than YAIM) is almost complete for usage and now I have a better understanding of the development cycle of the individual components; I am less worried about the move away from a single product release.



11 December 2013

Off to Edinburgh for DPM Workshop 2013

Friday has me going to Edinburgh for the latest DPM workshop. It'll be good to meet and discuss all the new(ish) features DPM has to offer (and make my own suggestions...)

Meeting agenda should be available at the following indico meeting page:

http://indico.cern.ch/conferenceDisplay.py?confId=273864

22 October 2013

DPM Load Issue Mitigation for Sites seeing load brownouts on individual disk servers.

For a while, it has been known that sites can experience load issues on their storage, especially for those sites running a lot of ATLAS analysis.

For DPM sites, this is caused by a combination of poor dataset-level file distribution (something which a storage system cannot guarantee without knowledge inaccessible to it apriori, especially with the new Rucio naming scheme), and the lack of any way to consistently throttle access rates, or the number of active transfers, per disk server.

At Glasgow, we found the following mixed approaches seemed to significantly ameliorate the issue:


  • Firstly, enabling xrootd support in DPM (this is now entirely doable via YAIM or Puppet; and it's required for ATLAS and CMS xrootd federation to work anyway), and setting the queues in AGIS to use xrootd direct io (rather than rfcp or xrdcp) for remote file access. A majority of Analysis jobs access only fractions of the files they rely on, so this reduces the total amount of data that a job needs to move, and also distributes the load caused by the data accesses over a longer time period, reducing the peak IOs required by the disk server. Getting your queue settings changed requires you to have a friendly ATLAS person with relevant permissions available.
  • Secondly, changing queue settings in your batch system to limit the rate at which analysis pilots can start. The cause of IO brownouts on disk servers seems to be the simultaneous start of a large number of analysis jobs (all of which immediately attempt to get/access files from the local SE); limiting the rate of pilot starts smears this peak IO over a longer time, again smoothing load out.

    With our torque/maui system, we accomplish this by setting the MAXIPROC value for the atlaspil group to a small number (10 in our case). MAXIPROC sets the maximum number of jobs in a class which can be eligible for starting in any given scheduling iteration - essentially it means that maui will not start more than 10 atlaspil jobs every (60 second) scheduling iteration, in our case.
    i.e:           GROUPCFG[atlaspil] {elided other config variables} MAXIPROC=10

With these two changes, we rarely see issues with load spiking and brownouts, despite increasing our maximum fraction of Analysis jobs from ATLAS significantly since the fixes were installed. Evidence suggests that both changes are needed for robust effects on your load, however.

30 June 2013

Big data...

Big data is a big buzzword these days but it's also a real "problem" (in some practical sense: data can be difficult to move, to store, to process, to preserve - when you have lots of it), but the big data is also an opportunity.

Well, we've had our big data workshop. A proper writeup will be, er, written up, but here are a few personal notes to whet your appetite (hopefully). You can of course also go and view the presentations at the workshop - all speakers have made their presentations available.
  • Lots of research communities have "big data" - in fact, it's hard to think of one that doesn't. Perhaps it's like email - after email became widely used at universities, researchers could communicate rapidly with each other, thus extending remote collaborations - but it's also a double edged sword in the sense that we spend much of our time processing email. Big data tools, I think, enable each research community to do more, but perhaps at a cost.
  • The cost is not always obvious beforehand - many speakers mentioned the complexity of software, and maintaining software for your data which makes use of hardware advances - and, er, runs correctly - is nontrivial.
  • Likewise, adapting existing code to process "big data" - to which extent is it like adapting applications and algorithms to run parallel, distributed, in the grid, in the cloud - ?
  • There is clearly opportunities for sharing ideas and innovations - the LHC grid may be finding events not inconsistent with the existence of Higgs-like particles (or, to the less careful, discovering the Higgs), thanks to a global network of data transfer, storage, and computing (of which GridPP is a part) crunching data on the order of hundred(s) of petabytes. But this work doesn't rely on visualisation to the extent that astronomy does. And in humanities, artists have found new ways of visualising data.
  • And who mentioned software complexity - the compute evolved along with building the collider, so we had time to test it.The last thing a researcher wants is to debug the infrastructure (well, actually, a few quite like that, but most would rather just get on with the research they're supposed to be doing.)
  • Data policies - open data, sharing - making data sets usable - and giving academic credit to researchers for doing this. There is more to big data in research than the ability to store it. Human stuff, too. Policy. Security. Identity management. That sort of stuff.
  • I think there is a gap between hardware/tools on one side and communities on the other - and the bridge is the infrastrcuture provider. But there is sometimes more to it than that - some communities find it harder to share than others. A cultural change may be needed.
  • And making use of big data tools - as with clouds, grids - sometimes it's easier running stuff locally, if you can, or it feels more secure. Make use of tools that make it easy. Learn from others.
Anyway, these are quick thoughts - we will have a proper writeup soon...

06 June 2013

Demonstrating 100 Gb/s

One of the interesting things from a networkingdataological perspective at this year's TNC is showing 100 Gb/s link across the Atlantic, and also the conference itself is connected with 100 Gb. 100 Gb is here!

We also had a music performance with a local (Maastricht) band and one musician in Edinburgh, and they were playing together. This can of course only happen if the latency is very low - 10-20ms. First time I have seen this in practice, very impressive.

26 April 2013

My new sister Eve

SO I have a new friend Eve. Eve is my last sister born in 2012 whose first home is the same as mine.
Initial info of Eve to compare with mine  and Georgina  follows. What surprises me is the number of files that that have no replicas at all and so are at risk if a house or room gets destroyed.


DataSet Name Dave  Georgina Eve
"DNA" Number
2.3.7.3623 3.3.3.3.5.5.89 3.53.1361
Number of Countries
17 17
Number of "Houses"
59 79 63
Type of Rooms:DATADISK 28 49 50
Type of Rooms:LGD
37 61 14
Type of Rooms:PERF+PHYS
32 58 22
Type of Rooms:TAPE
6 12 9
Type of Rooms:USERDISK
1 9 5
Type of Rooms:CERN
3 6 5
Type of Rooms:SCRATCH
0 12 14
Type of Rooms:CALIB
0 5 7
Total number of people (including clones) 1166 1594 642
Number of unique people
894 1120 299
Numer of "people" of type:
^user 137 470 64
Numer of unique "people" of type:
^user 132 429 55
Numer of "people" of type:
^data 725 1048 538
Numer of unique "people" of type:
^data 532 631 212
Numer of "people" of type:
^group 31 70 40
Numer of unique "people" of type:
^group 26 54 32
Numer of "people" of type:
^valid 1 0 0
Numer of unique "people" of type: ^valid 1 0 0
 Datasets that have 1 copy 691 811 178
 Datasets that have 2 copies 143 197 63
 Datasets that have 3 copies 52 68 25
 Datasets that have 4 copies 71 23 7
 Datasets that have 5 copies 1 7 21
 Datasets that have 6 copies 0 5 2
 Datasets that have 7 copies 0 2 1
 Datasets that have 8 copies 0 1 1
 Datasets that have 12 copies 0 0 6
 Datasets that have 13 copies 0 0 3
Number of files that have  1 copy 53644 141663 9509
Number of files that have  2 copies 12904 36863 7705
Number of files that have  3 copies 1860 3708 4520
Number of files that have  4 copies 7 2092 110
Number of files that have  5 copies 7 572 1091
Number of files that have  6 copies 0 487 20
Number of files that have  7 copies 0 142 5
Number of files that have  8 copies 0 73 1
Number of files that have  12 copies 0 0 6
Number of files that have  13 copies 0 0 5
Total number of files on the grid
85095 242241 22973
Total number of unique files on the grid 68422 185600 44674
Data Volume (TB) on the grid that has  1 copy 8.8 27.7 1.7
Data Volume (TB) on the grid that has  2 copies 6.8 21.3 4.7
Data Volume (TB) on the grid that has  3 copies 0.3 5.6 7.2
Data Volume (TB) on the grid that has  4 copies 0.011 1.5 0.12
Data Volume (TB) on the grid that has  5 copies 0.028 1.4 0.47
Data Volume (TB) on the grid that has  6 copies 0 0.36 0.039
Data Volume (TB) on the grid that has  7 copies 0 0.13 < 1GB
Data Volume (TB) on the grid that has  8 copies 0 0.12 < 1GB
Data Volume (TB) on the grid that has  12 copies 0 0 < 1GB
Data Volume (TB) on the grid that has  13 copies 0 0 < 1GB
Total Volume of data on the grid (TB)
23.7 104 35.8
Total Volume of unique data on the grid (TB) 16 58 14.2