Skip to main content

Week 2

Week 2

Open Data

Open data, by definition, is data that is free to use by anyone for any purpose at any time. This includes republishing data and not being affected by Copyrights or Patents by doing so. One proposed example of this is satellite data. By allowing this data to be "open data" we can allow it to be analysed to learn more about how our atmosphere is being affected by greenhouse gasses as well as monitor current changes and prevent new problems from arising. 

https://www.futurelearn.com/courses/big-data-and-the-environment/3/steps/420374


Data Scientist

A data scientist is someone who's job it is to analyse and interpret data, like usage statistics for a website in order to aid/assist a company or business with making important decisions. IEA data scientist Dr Ben Lloyd-Hughes states that there are 7 phases of a data science project which a data scientist will complete.


1. Problem Statement

2. Data Acquisition

3. Data Preparation

4. Data Exploration

5. Model Building

6. Documentation

7. Publicity Material

https://www.futurelearn.com/courses/big-data-and-the-environment/3/steps/420379

Metadata

Metadata can be defined as "data describing data". It can include a wide range of information including who wrote the data, why it was recorded, when it was recorded, the units of measurement the data is in, and if any copyrights are contained on the data. Metadata can be vital when wanting to compare on data set with another, especially when the data has different authors.

https://www.futurelearn.com/courses/big-data-and-the-environment/3/steps/443652







Comments

Popular posts from this blog

7. Limitations of traditional data analysis

7. Limitations of traditional data analysis Security: Security is foremost aspect for every technology. Big data is prone to data breaches. The important information that is provided to some third party may get leaked to customers. Proper encryption must be made in order to protect the data.  Large growth in data : data is growing faster than the processing power. Large volumes of data are being exploded in past years. We need some new machines to work; otherwise we will get over run by data. Large Data centres can solve this problem.  Inconsistencies: Sometimes the tools we use to gather big data sets are imprecise. This will happen when the data is collecting for example, consider Google search Edinburgh College results of the search on one day will be different from other day, this is mainly due to inconsistency in data collection.

Week 3

Week 3 Big and small data Small data is a way of keeping the size of data down. A file such as an image could take a lot of space depending on the size of the image, but by compressing it and losing some of the quality you can massivly reduce ths size. An example of small data could be Meteorological Aviation Report (METAR). The image below shows this off, each section of the code represents a different piece of information and saves entire data entries from being typed out and takes less time to process.  https://www.futurelearn.com/courses/big-data-and-the-environment/3/steps/420387 Citizen Science Is the collection of data about the natural world that has been gathered by the general public, usually as part of a project with other data scientists. This can be used to gain vast amounts of information quickly since many people at once can gather the information as opposed to just one scientist or one sensor. One example is Thames 21, a charity which worked with Cit...

2. Historical development of Big Data

2. Historical development of Big Data 1980's - networks started to appear across other countries, although possible, connection was still extremely difficult without travelling, the networks had to communicate in the same language as opposed to different ones. The principle link at CERN connected the US and Europe in 1989 1990's - remote access of data in the terabytes was now easily achievable around the world. To share data more easily, the web was created so that the actually location of the data wasn't necessary to know 2000's - Big data was now so large that CERN was unable to store such a vast amount of petabytes themselves and as such had to share it between partners around the world to offload some of it. This lead to the creation of cloud computing 2005 - The term Big Data is used for the first time by Roger Mougalas, a year after the creation of the term "web 2.0" which refers to data too large to handle with traditional business methods 2009 -...