Sunday, 1 March 2015

A G-Cloud Thought Experiment...

For those of you that may not be familiar with the G-Cloud, it's the procurement framework set up to enable UK Government bodies to procure access to cloud services.   It's not a cloud service itself, purely a procurement framework.   I've taken a keen interest in G-Cloud over the years for a number of reasons:

i) I like cloud computing so much I wrote a book on securing cloud services;
ii) I worked client-side at a Government agency helping them to move all of it's IT delivery over to G-Cloud services; and
ii) I was the Industry security subject matter expert on the original Intellect/CIO Council project team that first floated the idea of a Government cloud (the HMG Data Centre Consolidation project).  I was there at the start.

The reason for resuscitating my blog is the changing approach towards assuring cloud services. 

In the earlier iteration of the G-Cloud framework, cloud services were assured by the Pan-Government Accreditors prior to being deployed by individual departments.   This is now changing. G-Cloud providers can now simply assert their security credentials and it's up to each G-Cloud customer to decide whether this is sufficient to meet their needs.   This is plainly silly.   We've moved from a single body providing a consistent level of assurance to all of Government - do it once, do it right - to an approach that requires *each* department to perform it's own assurance.  

If the centre was struggling to hire/train enough resource to assure the G-Cloud services, I'm wondering where the skilled resources will come from to staff each of the individual HMG bodies that now need to conduct their own assurance activities?  Clearly it's not going to happen.    I haven't even touched on the folly of relying on self-assertion of security controls by service providers.   This is similar to boarding a Trans-Atlantic flight and hearing the pilot on the tannoy stating "Don't worry folks, we're not doing any safety checks before take-off as I know you're keen to get going.  But we will of course do some checks when we land.   Or otherwise cease to be in the air...".

But this is a post about a thought experiment.   So, here we go...

Department A has traditionally relied upon Pan-Government Accreditation to assure cloud services before they adopt them.   Given the change in approach, the Department has now decided that it won't rely simply upon self-assertion but will also require any cloud provider to hold an appropriately scoped ISAE3402/SSAE16 SOC 2 or CSA STAR certification.    This ruled out Cloud Provider X which only had the self-asserted statements.

Department B has a higher appetite for risk than Department A.  They are happy to accept self-assertion.   They've therefore adopted Cloud Provider X because it offers extremely cheap services (for some reason).

Department A needs to share data with Department B.    Department A knows that Department B makes use of Cloud Provider X - a cloud provider that Department A has explicitly already decided is not sufficiently assured to host their data.

What happens now?   Does Department A go explicitly against it's own local risk management and share the data regardless?   Does Department A share the data with Department B alongside a "common sense" "plain English" set of handling conditions that state that the data must not be processed on Cloud Provider X?   If so, where does that leave Department B?   With a need to have an alternative to Cloud Provider X or simply ignore the handling conditions, that's where.   Not a good place to be.

This could all have been avoided by having a more sensible approach to cloud provider assurance, one which encouraged a common level of assurance and so facilitated data sharing.  

Unfortunately, HMG security policy is moving away from a usable set of common taxonomies and standards and towards a more flexible, locally risk-managed approach which sounds great in theory but is destined to fail horribly in practice as departments each go their own way, adopting different interpretations of "commercial best practice" and making data sharing much, much harder...





Friday, 28 March 2014

Full Disclosure is dead. Long live Full Disclosure!

So, earlier this month the long established Full Disclosure mailing list was closed – or at least suspended indefinitely after the list moderator (John Cartwright) finally reached the end of his tether after many years of legal threats and gratuitous trolling.   The e-mail sent to the mailing list outlining the reasons for it’s suspension can be accessed from the below link:
Now despite it’s faults, including the low signal to noise ratio, the Full Disclosure list acted not only as a venue for genuine full disclosure of security vulnerabilities but also as a venue for some of the more “interesting” aspects of hacker culture to be displayed.   After a little wailing and gnashing of teeth in the security community (and some discussion of replacing the mailing list with the #fulldisclosure Twitter hashtag), Fyodor (aka Gordon Lyon, creator and maintainer of nmap) has now revived the mailing list with the blessing of the moderator of the original incarnation.   Fyodor’s note announcing the revival of the list can be found at:
Now, my reason for writing this post is not just to highlight the welcome return of the Full Disclosure mailing list (thanks Fyodor) but more to highlight the vulnerability of some of our more treasured Infosec resources.   I think there’s a tendency to forget that resources such as Full Disclosure are not pain-free for those that provide them.   In some cases there’s a commercial reward for providing the resource, e.g. vendor mailing lists.   But where those resources are provided out of a sense of community, I think we need to be more appreciative of the humans behind those services.   Sure, trolling can be amusing for those with low/weird entertainment thresholds, but there’s no point taking it to the extreme such that you lose one of your avenues for displaying your self-perceived awesomeness.   Similarly, if you post something to an e-mail list called Full Disclosure, please don’t be surprised if it remains on the Internet alongside the responses of those that don’t agree with you.   If you’re not happy for your opinion to be eternally available then don’t post it to the Internet.   Certainly don’t start throwing legalese at those providing the community resource otherwise or they may just take it away.  
I’m certainly not advocating an absence of trolling, flaming or arguing on the re-born Full Disclosure as that would go against the established culture of the list but it would be good if the legal threats could be left behind with the old iteration…

Saturday, 22 February 2014

My thoughts on Care.data

Well, I didn't expect to be reviving this personal blog however I needed to make it clear that this post is absolutely just my personal opinion.


I've been involved in a Twitter discussion with @dbarthjones and it's become apparent I needed more than 140 characters to adequately express my views with respect to the NHS data sharing initiative known as care.data.


In my opinion, Ben Goldacre hits the nail on the head in his post at:


http://www.theguardian.com/society/2014/feb/21/nhs-plan-share-medical-data-save-lives


In particular he notes that the public have two main concerns with the proposals:


i) Privacy; and
ii) Commercial exploitation of their personal data.


Jumping straight to sharing of personal data with the commercial sector was a clear presentational mistake if nothing else.   But I'm not blogging about (ii), I'm more concerned with (i).


I'm not overly concerned about the potential general re-identification of all individuals within the care.data dataset - at least not at present.  There are few datasets out there of sufficient scale and cross-over (shared attributes) to make general re-identification trivial and easy to automate.   Of course, this will change if the Open Data evangelists get their way and add education, tax and other HMG assets to the pot.  In this latter scenario the kind of general re-identification across these huge data-sets (complete with shared unique sets of attributes) becomes much more likely/scary.   No, what I am concerned about today is the targeted re-identification of individuals by those who can have a real impact on our day to day lives; think friends, family, employers etc.  Such potential threat actors already know a great deal of information about us - information that we have chosen to share with them.   That does not mean that they know everything about us, we are perfectly entitled to keep things to ourselves.   If we fall ill, we are not obliged to tell our friends exactly what our condition is or even to tell them that we are ill at all!    And that is what worries me, if our data falls into the public domain then we lose that ability to decide with whom we share our most private information.    As far as I am aware there no current plans to make individual-level care.data publicly available, only to researchers, however the more widely you distribute data the less control you have over further distribution.  It comes down to trust in those with whom you share, albeit trust usually bound up in some of data sharing agreement.


Now, there are lots of bland statements made by proponents of data sharing that they have been sharing health data for years and that there have been no major incidents of data leakage.  My question is this:  how do they know?  What monitoring systems are in place?  You can be very good at not spotting security breaches if you don't look.  Having spent many years working across HMG I am more than aware of the amount of spare resource that the Civil Service currently has available to police enforcement of data sharing agreements.


So.  That's the bad stuff.  I do however fully buy-in to the potential benefits of sharing health data.  Conceptually, I know it makes sense and, to be honest, prior to the last couple of weeks I was not planning to opt out.   My opinion has been hardening due to the utter failure of those pushing the scheme to acknowledge genuine concerns.  See http://www.bbc.co.uk/news/health-26277866 for example.  If they either don't understand, or choose not to address, concerns around re-identification and commercial exploitation do I really want to trust them with my data?   I have not yet opted out.  I want to be convinced that the people running the scheme understand the issues and are willing to seek our consent for this sharing of OUR data.  At the moment I fear it is still a case of public opinion being ridden over rough-shod in favour of commercial interests.   That would be a shame.   There is a real opportunity to sell the benefits of Open Data but this MUST be done in a balanced banner and which gives data subjects genuine choice.   It's a question of benefit and risk but we should be able to make our own individual choices and not have the decisions made on our behalf by those who think they know best.  They don't.

Wednesday, 22 May 2013

Anonywhat?


 
There are very few areas that highlight the fundamental conflict between security and usability better than data anonymisation.  It’s a conflict that’s only going to get worse as more and more data is collected about us all as individuals.  A recent blog post by Bruce Schneier expresses his concerns over the Internet of Things and the vast increase of personal data available for collection once our cars, fridges, televisions, smart meters et al are all on the Internet and ‘measuring’ our usage for reporting back to their vendors or service providers.  Unfortunately it’s the onward step from there where the vendors and service providers sell on that data to ‘selected partners’ that the real loss of privacy, and gain of unsolicited opportunity to partake in exclusive new offers, starts to kick in.   It wouldn’t be so bad if it was only the corporate types looking to pimp details of our personal activities and opinions to all and sundry; unfortunately we also have the UK government looking to ‘maximise the value’ of the data that it holds about us.  For example, by releasing individual records of pupil attainment to enable industry to release it’s full vigour and expertise and create lots of wonderful new tools and products to improve education in the UK.  Just don’t ask to see the actual business case – it’s so obvious that this is the way it would work that it’s simply not worth asking the question…   But don’t worry, all such data, be that sourced from the Internet of Things or HMG databases that we have no choice but to populate, will be anonymised before being released.   And that’s when the problem hits…
 
First off, let me start by explaining what I mean by anonymisation, and it’s counter-process de-anonymisation.   The process of anonymisation is supposed to render it infeasible for an attacker to identify a real-world identity from a dataset.  Conversely, de-anonymisation is the process of identifying a real-world identity from anonymised data.   For the purposes of this blog I’m not going to explore the concept of pseudonymisation as it’s not directly relevant – but feel free to Google the term (irony).
 
So, how can you safely anonymise data?  There are a number of techniques available, from the highly noddy (removal of obviously personally identifying information such as names and addresses) through to the more mathematically valid (but still imperfect) approaches such as k-anonymisation, t-closeness and l-diversity.  Most organisations seem to fall somewhere in the middle and use techniques such as data substitution, data perturbation, data aggregation and data suppression.   Now, the best way to safely anonymise data (imho) is to aggregate data and suppress small numbers.  This means that you don’t actually release data on individuals but rather release data on groups and suppress (i.e. strip out) data that would identify small groups of people (e.g. 5 or less).   Good examples of aggregated data include the school performance tables that provide details of exam pass-rates.  However, data aggregates are not good for analysis of individuals – and that makes it of little use for those organisations looking to pitch to those customers with most interest in their products or services.   The naïve will think that simply stripping out name and address and other obviously identifying information will make their data safe for re-use.  This is demonstrably wrong as illustrated by the cases of AOL, NetFlix and the more recent research on the uniqueness of mobile telephony data.   Let’s take this last case as an example of the issue.  The study in question showed that it can only take 4 geographical data points within the mobile telephony dataset to identity an individual.   There are no names in the data, no other obvious means of identification, just location data.  I’ll give you an example of how that’s a privacy risk.  I live near one major city but my current client is based in another, some 2.5 hours away by train.  This means that every day there will be at least three data points regarding my mobile phone usage – my home address, my local Train Station and my client office.  Tie this together with my regular visits to my Taekwondo classes and that’s pretty much me identified – I’m the only one visiting the location of my Taekwondo class that also regularly visits my client offices!   Of course, in order to resolve my identity in this way, you need to know my work patterns and that I do Taekwondo – nothing that a nosy neighbour or friend on Facebook would find difficult to discover.   The privacy risk at this point is that, once they have identified my unique identifier (either a genuine unique identifier or simply a unique recurring relationship of elements within the data) within the dataset they can then start to increase the information that they know about me – for example, they may now spot the football ground that I attend or, more worryingly, should I be ill and start visiting the hospital on a regular basis this would also become apparent to those with whom I have no wish to share that information…
 
This takes me back to the fundamental premise of this blog – the conflict between security and usability.  The only way to truly anonymise data is to destroy the relationships within the data that enable a knowledgeable attacker to identify the individual.  There are two problems with this:
 
i)                    you will rarely know the exact set of data, and therefore relationships, known to an attacker.  Consider the ‘nosy neighbour’ threat actor; my own neighbours know our names, address, birthdays, the cars that we drive, the names of our kids and the school(s) that they go to.  Furthermore, they know our hobbies and the hobbies of our kids, they know where we grew up and who we work for…  That’s an awful lot of useful information to discriminate between individuals within a dataset [by the way, we have lovely neighbours and so, personally, I’m not too worried.]
ii)                  it’s the relationships that give the data value!  If you perturb, substitute or otherwise mangle the data then you risk losing the relationships in which you are most interested.  For example, if the data collector is a retail outlet and they mangle their dataset by including random purchases from other shoppers within my loyalty card history (so as to make it less obviously me) then they risk starting to target me with offers that I have absolutely no interest in and driving me towards alternative retailers.
 
Now, there is guidance available (e.g. the ICO document entitled Anonymisation: managing data protection risk code of practice) and there are tools in the market to help organisations to anonymise their data, however you really need to understand what you are doing to get the most benefit out of such tools.   More important than the how though is the why.  Organisations should ensure that they have a rock solid business case outlining genuine, evidenced, benefits to both themselves and the data subjects before releasing anonymised datasets to the public or to their business partners.   A failure to develop such a strong business case will leave organisations highly exposed should their anonymised dataset not be as anonymous as they thought – the data subjects may not be amused to find that their intimate personal details have been made available to all simply because someone thought it may… possibly… potentially… be useful.  Remember this, once the data has gone, the data has gone; there is no way to put this particular genie back in it’s bottle.
 
 
 
 
 

 

 

 

 

 

 

Thursday, 5 July 2012

Security Architecture - it's on video now...

I've been blathering about the benefits of adopting architectural approaches to security for a while now.  I was recently asked to appear in a video for Professional Outsourcing magazine to put forward some views on security architecture and the benefits that adopting such an approach can deliver.   Here you go:

http://professionaloutsourcingmagazine.net/videos/security-getting-in-way-your-business

Feel free to share your thoughts :-)

Thursday, 24 May 2012

Too cheesy?

I came up with the below analogy for a presentation I gave earlier this week... too cheesy?
Adopting the cloud is like learning to dive - you need to get your technique right upon entry and look out for the sharp and pointy things lurking under the surface. Get it right and you'll improve your flexibility and agility, get it wrong and you could well break something.

 Of course, the only way to finish off is "Look before you leap!".   Just thought I'd share :-)

Friday, 4 November 2011

Evidence-based opinions

I came across a couple of interesting web-sites over the last couple of weeks that I think are worth sharing. The first of these relates to work conducted by the Australian governments Defence Signals Directorate (DSD). Through analysis of the vulnerabilities and exploit attempts reported to them, the DSD has drawn up a set of 35 mitigations that would have helped to prevent exploitation. In fact, just implementing the top 4 strategies would have "prevented at least 70% of the intrusions that DSD analysed and responded to in 2009, and at least 85% of the intrusions responded to in 2010". What are these top four strategies?

 patching third party applications;
 patching operating systems;
 minimising administrative privileges; and
 application whitelisting.

The first 3 should be just good practice. The 4th one can be more difficult to get past by the business. In any case, it's nice to see a set of mitigation strategies based off real analysis rather than simple reliance on 'best practice'. The DSD documents can be found over at:

http://www.dsd.gov.au/infosec/top-mitigations/top35mitigationstrategies-list.htm
http://www.dsd.gov.au/infosec/top35mitigationstrategies.htm

Definitely worth a read.

What else has caught my eye? Well, I have no choice but to give a shout-out to the competition. PwC have released their 2012 Global State of Information Security Survey and have provided a nifty way of exploring the underlying data - available over at:

http://www.pwc.com/gx/en/information-security-survey/giss.jhtml

As ever, the GISS is a worthwhile read and the highlights for me relate to the cloud security aspects:

 It's tight, but there are now more respondants saying that they use cloud services than there are saying that they do not. The Do Not Knows could still tip the balance either way though!
 SaaS still holds a hefty lead as the most commonly implemented service model, followed by IaaS and then PaaS.
 Of those who have implemented cloud services, over half believe that the move to cloud has improved their security. Less than a quarter believe that the move has weakened their security.

I've been blathering on for a while that a move towards cloud services can have security benefits as well as the more often documented downsides. It's re-assuring to see that a majority of those moving towards the cloud believe that the positives actually outweigh the negatives.