09 March 2012

Busy afternoon for ATLAS Storage at the RAL T1

Lunchtime on the 8th March 2012 was a busy time for the disk servers in the ATLAS instance of CASTOR at the Tier1 at RAL. Servers were delivering approximately 45Gbps as a source:



FTS controlled transfers at this as controlled by the RAL FTS WAN transfers were negligible; since the majority of "WAN" transfer via FTS (RAL) system were actually an intra-CASTOR transfer.

"WAN Rate"


"Intra-CASTOR" rate shown below shows that of the 4-5GB/s only 2.5Gbps is internal FTS traffic. ( ~5% of traffic.)

However this is only ~1/3 of the outbound traffic over the WAN. The other traffic is controlled by other FTS servers.
ATLAS aggregate this into the following ATLAS specific plot:
This shows the outbound traffic was actually 625MB/s (~5Gbps)
 ATLAS have a graph showing amount of data processed . This peaked at over 8TB in one hour.


However 8TB/hour is approximately 25Gbps
So what was using the remaining bandwidth...

Well we also need to include traffic from disk servers to tape servers (which are LAN transfers that neither ATLAS nor FTS know about. This rate can be seen here:


This shows there is an extra 600MB/s (~5Gbps) of rate unaccounted for. With FTS+WN+Tape network loads; the known bandwidth usage is 35Gbps. As a rough estimation, this explains how the majority of the bandwidth was used.

08 March 2012

I have been quiet for a while so thought I just give a sit-rep regarding my numbers. I am in 91 houses; 283 rooms in total. The houses (and the number of rooms) in which I and my children currently stay are:
1 CA-MCGILL-CLUMEQ-T2
1 CSTCDIE
1 CYFRONET-LCG2
1 IL-TAU-HEP
1 INFN-ROMA3
1 INFN-TRIESTE
1 MAIGRID
1 NCG-INGRID-PT
1 RO-02-NIPNE
1 RO-07-NIPNE
1 RRC-KI
1 SMU
1 TUDRESDEN-ZIH
1 UKI-LT2-UCL-HEP
1 UKI-SOUTHGRID-CAM-HEP
1 UNIBE-LHEP
1 UNI-BONN
1 UNICPH-NBI
1 UNIGE-DPNC
1 UPENN
1 VICTORIA-LCG2
1 WISC
2 AUSTRALIA-ATLAS
2 BEIJING-LCG2
2 CA-ALBERTA-WESTGRID-T2
2 CSCS-LCG2
2 ILLINOISHEP
2 IN2P3-CPPM
2 INFN-BOLOGNA-T3
2 INFN-FRASCATI
2 LIP-COIMBRA
2 LIP-LISBON
2 NERSC
2 OU
2 RU-PROTVINO-IHEP
2 SE-SNIC-T2
2 TR-10-ULAKBIM
2 UAM-LCG2
2 UKI-NORTHGRID-SHEF-HEP
2 UKI-SOUTHGRID-BHAM-HEP
3 GOEGRID
3 GRIF-LPNHE
3 IFAE
3 IN2P3-LPC
3 IN2P3-LPSC
3 JINR-LCG2
3 PIC
3 SARA-MATRIX
3 SFU-LCG2
3 UKI-LT2-RHUL
3 UKI-NORTHGRID-MAN-HEP
3 UKI-SCOTGRID-ECDF
3 UKI-SOUTHGRID-OX-HEP
3 WEIZMANN-LCG2
4 AGLT2
4 CA-SCINET-T2
4 CA-VICTORIA-WESTGRID-T2
4 FZK-LCG2
4 GRIF-IRFU
4 GRIF-LAL
4 IFIC-LCG2
4 IN2P3-LAPP
4 INFN-T1
4 LRZ-LMU
4 MPPMU
4 MWT2
4 NET2
4 PRAGUELCG2
4 RAL-LCG2
4 SWT2
4 UKI-LT2-QMUL
4 UKI-NORTHGRID-LANCS-HEP
4 UKI-NORTHGRID-LIV-HEP
4 UKI-SCOTGRID-GLASGOW
4 UKI-SOUTHGRID-RALPP
4 UNI-FREIBURG
4 WUPPERTALPROD
5 DESY-ZN
5 INFN-MILANO-ATLASC
5 INFN-NAPOLI-ATLAS
5 INFN-ROMA1
5 TAIWAN-LCG2
6 DESY-HH
6 SLACXRD
6 TRIUMF-LCG2
7 NIKHEF-ELPROD
7 TOKYO-LCG2
8 BNL-OSG2
8 IN2P3-CC
8 NDGF-T1
10 CERN-PROD

The type (and number) of rooms of each type in which my children and I live:
1 DATAPREP
1 DET-LARG
1 DET-MUON
1 PERF-MUONS
1 PPSDATADISK
1 TMPLOCALGROUPDISK
1 TZERO
2 CALIBDISK
3 PERF-EGAMMA
4 PERF-JETS
4 PERF-TAU
5 PERF-FLAVTAG
5 PHYS-EXOTICS
5 TRIG-DAQ
5 USERDISK
6 PHYS-BEAUTY
7 PHYS-HIGGS
8 DATATAPE
9 PHYS-SUSY
9 PHYS-TOP
11 PHYS-SM
55 SCRATCHDISK
56 LOCALGROUPDISK
71 DATADISK

As you can see the type and number of each type of room vary quite alot.

I have 1445 unique children (but over 2700 in total including clones.) Some of my children have up to 16 clones, worringly 989 of my children have no clones at all , should any on my houses or rooms become destroyed many may be lost. Since I am the unique ancestor of all of them, all are reproducable (with various amounts of effort.)

287/636 of the unique Dirks have no clones and 43/57 of the unique Gavins have no clones. These are the easiest to reproduce; as they are produced en mass centrally. However, 659/753 of the unique Ursalas have no clones and are the highest risk children to be able to be reproduced. If only we could get Ursalas' avatar owners to create more clones of them.

In total; my children are 218863 files and 47.065TB of data; whereas I was 15200 files and ~16TB. I.e my children are ~3 times my size; but 14 times my number of files. Of my children; 42TB (143k files) are Dirk's; 1TB(37k files) are Gavin's. Ursula's make up the remaining 5TB/ 3k Files.
My birthday calendar looks like:
Day Apr May Jun Jul Aug Sep Oct Nov Dec Jan Feb Mar
1
9
1 1
1 5 4 3 6 16
2
1 31
2
1 2 3 3 7 4
3
3 5 1

4
1 2 4 16
4

21 1

2 6 2 6 7 7
5
2



3 7 2 3 7 2
6
2
1
2 10 1 5 5 19 13
7




2 1 1 3 1 8 7
8




3
5 5
9
9
1
12
5

3 2 18
10
1 4

2
2 5 2 11
11
1 1 1 2 4
1 8 22 9
12
5
5 1 2 2 3 10 5 3
13
4
6 2 1 4 4 4 6 5
14
6 2 1

3 1 7 6 6
15

3 5 1 5 2 1 3 3 6
16
1 3
4 4 2 2 7 2 5
17

3 3 17 4 1

2 9
18
1 1 2 27 1 1
2 7 11
19



37
1
8 8 7
20
2

7

3 5
18
21
2

6
3 4 1 5 13
22

1
17 5 2 3 2 6 17
23
1 1 8 7 11 1 3 7 10 13
24

1 6 8 3 1 5 2 6 21
25 30 2 4 2 10 2 3 9
5 20
26
3
2 14 6 4 10 2 6 13
27 13 1
1 14 3 5 2 1 3 14
28 36

1 1
7 1
5 7
29 62 14 1 2
2 5 4 9 7 9
30 6
12 1

4 5 2 11

31


4

1
2 9


So you can see there was a busy period around my birth; then in the second half of August; and relatively busy again in the last six weeks.

26 February 2012

Reaganomics

It is very impressive actually - Reagan Moore running an iRODS tutorial session at the pre-ISGC2012 iRODS workshop and getting a whole room to install clients (and a few servers) and downloading files from RENCI in North Carolina. From the "other" grid perspective (ie SRM), ASGC's Hsin-Wei Wu presented work on developing an SRM interface to iRODS, like they previously developed it for SRB. We need to start interoperating these implementations and see what an information provider will look like - previously we had just static information providers for SRB, maybe we can do dynamic ones for iRODS.

Reagan asked the question whether there is an SRM backend for iRODS -- ie you have iRODS "on top" fetching data from an SRM storage element. This has, of course, been talked about before but there has never been a strong business case. Maybe now in the brave new world of multi- and interdisciplinary research, the case is stronger.

On the whole, it has been a very interesting tutorial - any tutorial which makes you want to go away and write code is surely a successful one :-)

23 November 2011

The best rate to get from ATLAS's SONAR test involving RAL can be assumed to be internal transfers from one Space token at RAL to another space token at RAL. The Sonar plot for large files; (over 1 GB,) for the last six months is:

Averaging this leads to:

Leading to average of 18.4MB/s as the average rate with spikes in 12 hour average to above 80MB/s. (Individual file transmission rates across the network (excluding overhead) have been seen at over 110MB/s. This relates well to the 1Gbps NIC limit on the disk servers in question.

Now we know that of Storm,dCache,DPM and Castor systems within the UK that Castor tends to have the longest interaction overhead for transfers. Overhead for RAL-RAL transfer varies for the last week is between 14 and 196 seconds with an average of 47 seconds and a standard deviation of 24 seconds.

12 November 2011

Storage is popular

Storage is popular: why, only this morning GridPP storage received an offer of marriage from a woman from Belarus (via our generic contact list). I imagine they will stick wheels on the rack of disk servers so they can push it down the aisle. We need a health and safety risk assessment. Do they have doorsteps in churches? Do they have power near the altar or should we bring an extension?  If they have raised floors, can we lay the cables under the floor And what about cooling?

Back to our more normal storage management, it is worth noting that our friends in WLCG have kicked off a TEG working group on storage. TEG, since you ask, means Technical Evolution Group - the evolution being presumably the way to move forward without rocking the boat too much, ie. without disrupting services. The groups role is to look at the current state, successes and issues, and how to then move forward - looking ahead about five years.  In good and very capable hands with chairs Daniele Bonacorsi from INFN and our very own Wahid Bhimji from Edinburgh, the group membership is notable for being inclusive in the sense of having WLCG experiments, sites, middleware providers, and storage admins involved. Although the work focuses on the needs of WLCG, it will also be interesting to  compare with some of the wider data management activities.

11 November 2011

RAL T1 copes with ATLAS spike of transfers.

Following recent issues at the RAL T1 , we were worried about not just overall load on our SRM caused by ATLAS using the RAL FTS, but also the rate at which they put load on the system.
At ~10pm on the 10th November 2011 (UTC); ATLAS went from running almost empty to almost full on FTS channels involving RAL being controlled by the RAL FTS server. This can be seen in the number of active transfer plot:

This was caused by atlas suddenly putting into the ATLAS FTS many transfers which can be seen in the "Ready" queue:

This lead to a high transfer rate as shown here:
And is also seen in our own internal network monitoring:

The FTS rate is for transfers only going through the RAL FTS. ( I.e does not include puts by CERN FTS, Gets from other T1s or the chaotic background of dq2-gets, dq2-puts and lcg-cps not covered in these plots. Hopefully this means our current FTS settings can cope with start of these ATLAS data transfer spikes. We have seen from previous backlogs that these large spikes lead to a temporary backlog ( for a typical size of spike;) which clears well within a day.

25 October 2011

"Georgina's " Travels

So I received various postcards from Georgina's family from their new homes around the world.
~9 months after Georgina's birth she has:
1497 unique children in 79 Houses in a total of 265 rooms.
( Further analysis is hard to describe in an anthropomorphous world since data set replicas would have to involve cloning in Dave and Georgina's world.)
Taking this into account of the 1497 datasets , the distribution of number of replicas is as follows:
556/1497 datasets only have one copy.
Maximum number of copies of any "data" dataset is 20.
Maximum number of copies of any "group" dataset is 3.
Maximum number of copies of any "user" dataset is 6.
What is of concern is to me is that 279/991 user or group derived datasets have one unique copy on the grid.

12 October 2011

Rumours of my storage have been somewhat exaggerated?

It has been reported that RAL's tape capacity has grown by some factor, by which I deduce as the most likely explanation that at least one of the backend databases has been upgraded from 2.1.10-0 to 2.1.10-1:



/* Convert all kibibyte values in the database to byte values */
UPDATE vmgr_tape_denmap
   SET native_capacity = native_capacity * 1024;
UPDATE vmgr_tape_pool
   SET capacity = capacity * 1024;
As you can see the internal accounting numbers are multiplied with a factor 1024 which obviously confuses the CIP. The new CIP (2.2.0) has code to deal with this, but we can backport it to the current one. The caveat is that not all CASTOR instances may have been updated; we will check that.

27 September 2011

SRM speedup update

Or should that be speedupdate? If you remember the Jamboree last year, in Amsterdam, one of the suggestions to decrease the negotiation overheads in SRM by making more efficient use of the socket. Our very own Paul Millar from DESY has come up with a demonstrator using a lua shell which is able to reach SRMs by calling the functions in the API, like S2 does, but with perhaps a simpler language to learn, and you're in a shell.

What Paul demonstrated was the speedup associated with calling each function individually, and then by turning off GSI delegation, and finally by reusing the socket using HTTP KeepAlive. You'd not be surprised to see a big improvement - but of course the server must support KeepAlive.

Combined with the immediate return on srmGet when a file does not need staging, this could again speed up multiple file accesses (and of course you can still submit multiple file requests in a single SRM request.)

Paul has published the code, you can find the dCache LUA SRM interface.on the dCache web site.

14 September 2011

Bringing SeIUCCR to people

Am at the SeIUCCR (pronounced "succor" - no, not "sucker") summer school at the Coseners House in Abingdon and the doors are open out to the garden and we can well believe it is a summer school. Last time I lectured in a summer school (cloud security) I had made the presentation a bit too easy, so this time (data management) I included some hairy stuff. While it was basically about uploading data to the grid and moving it around, the presentation covered the NGS and GridPP, i.e. Globus and gLite, and we also (once) queried the information system directly (which was the aforementioned hairy part). But, like the 2-sphere, no talk can be hairy everywhere. Oh, and all the demos worked, despite being live.

The main idea is that grids extend the types of research people can do, because we enable managing and processing large volumes of data, so we are in a better position to cope with the famous "data deluge." Some people will be happy with the friendly front end in the NGS portal but we also demonstrated moving data from RAL to Glasgow (hooray for dteam) and to QMUL with, respectively, lcg-rep and FTS.

If you are a "normal" researcher (ie not a particle physicist :-)) you normally don't want to"waste time" learning grid data management, but the entry level tools are actually quite easy to get into, no worse than anything else you are using to move data. And the advanced tools are there if and when you eventually get to the stage where you need them, and not that hard to learn: a good way to get started is to go to GridPP and click the large friendly HELP button. NGS also has tutorials (and if you want more tutorials, let us know.)

It is worth mentioning that we like holding hands: one thing we have found in GridPP is that new users like to contact their local grid experts - which is also the point of having campus champions. We should have a study at the coming AHM. Makes it even easier to get started. You have no excuse. Resistance is futile.

05 September 2011

New GridPP DPM Tools release: now part of DPM.

I'm happy to announce the release of the next version of the GridPP DPM toolkit, which now includes some tools for complete integrity checking of a disk filesystem against the DPNS database.
This should also be able to checksum the files as well, although this takes a lot longer.

The bigger change is that the tools are now provided in the DPM repository, as the dpm-contrib-admintools package. Due to packaging constraints, this RPM installs the tools to /usr/bin, so male sure it is earlier in your path than the old /opt/lcg/bin path...

Richards would like to encourage other groups with useful DPM tools to contribute them to the repo.

04 August 2011

Et DONA ferentis?

If you've been reading papers in the past few years you would have seen DOIs in the references. (I mean academic papers, not newspapers.) The idea is that it saves you from writing

Journal of Theoretical and Applied Irrelevance, 53 Vol. 3 (1), (2008), pp.312-322.

when instead you can write a simpler string, a handle, which identifies the data.

For this to work, you will need a handle resolution service. Think of DNS. If you did it "by hand" you would resolve www.gridpp.ac.uk with:
dig www.gridpp.ac.uk. A IN
to an IP address, and then maybe telnet to port 80 or something. Or think of the GUID-to-SURL mapping in LFC that we all know and love (or the SURL-to-TURL mapping). Similarly, handles have to resolve into something that is the stuff you're looking for. Doesn't have to be a paper, it could also be data, or even particular versions of data.

Enter the Handle System. Patrice Lyons and Bob Kahn - one of the fathers of the Internet - from CNRI, are proposing to establish a global handle system, morally equivalent to ICANN, to manage the uniqueness and persistence of handles. This will be the Digital Object Numbering Authority, or DONA.

Of course, like DNS or GUIDs, there is no assurance that the data you're looking for is actually there. In fact, persistence is meaningful even for temporary objects, in the sense of the handle being associated with the object forever, even if the object itself doesn't live forever.

Sounds simple? Well, apart from those temporary objects, the system may need to be able to deal with modifiable objects, versions, replicas, and (possibly) part handles. Typing it in again from a printed representation. And what is the object? is it the object as a sequence-of-bits, or is it the "curation-object" which goes in as a Word97 document, say, and is referenced later as PDF.

We might even have a GFAL-type library which knows how to resolve handles into data, so the application doesn't have to know. Meanwhile, the handles are coming: apart from the publishers' DOIs in the papers, you can see the entertainment industry have also picked it up with EIDR.