Friday, February 4, 2011

Web Archiving Service Evaluation

Web Archiving Service Evaluation: "

This is the fourth installment in a series of evaluations of website harvesting software on the Practical E-records blog.  The first three installments were reviews of open source software that you can download and install locally—HTTrack, GNU Wget free utility, and Heritrix.  This fourth installment is a review of the Web Archiving Service (WAS) developed by the California Digital Library, which is a fee based service for capturing and storing websites.



The Web Archiving Service provides tools and support for harvesting websites and preserving them, as well as providing tools for analyzing the captured content.  For example, you can check between two captures to see how many web pages (and what pages) were changed or deleted.  Unlike the other web archiving software previously evaluated here, WAS is a fee based service for those outside the University of California (University of California organizations only pay for storage).  For those outside the University of California system, a yearly fee is required, but there are discounted rates for consortia of three or more institutions.  Thankfully, WAS also provides a free trial subscription to those institutions wishing to try out the software.


This is an excellent choice for those institutions that can afford a subscription.  It is very user friendly and produces a capture that has a high fidelity to the original site.  The largest concern I have is that the data is stored with WAS rather than on an internal server.  This might bring up long term preservation concerns if the subscriber is not able to commit to long-term support for the service or if CDL is not able to maintain the service for whatever reason.


Chris Prom contacted the WAS to get our trial subscription (on the right hand side of the page is info on how to contact them about getting a trial subscription).  We were then given a username and password to access the service over our internet browser.  I read over the WAS User Guide, skimmed some of the other documentation, and watched the user videos before beginning, but all of this is not necessary in order to begin.  Familiarizing yourself with the general system is important, but overall the service is straightforward enough that you could follow the manual as you are going through your first capture if you need to.


Once you’ve signed into the WAS site with your username and password, there are two main things you need to do to get the capture started.  You first need to create the site information and then capture the site.



Under the Create Site section you are able to adjust capture settings, scheduling, and add descriptive data.  Unlike the other web capturing software profiled previously, WAS has narrowed down many of the parameters that need to be adjusted for the capture.  This simplifies the process for those archivists who don’t have a lot of experience with more technical computer applications.  For example, with WAS, the capture settings that need to be adjusted are only scope, whether to capture linked pages, the maximum amount of time spent on the capture (1 hour or 36 hours), how frequent to make the capture (daily, weekly, monthly, custom), whether the capture will be made public, and added descriptive data about the site.  This is one of the many aspects of the WAS software that is well designed.  There are only a few easy to understand parameters to adjust, you can schedule the program to automatically capture the site on a regular basis, and you can add descriptive metadata associated with your capture.


After you have set up your site information you can click on the Capture Sites option which takes you to a “Manage Sites” page.  This page includes both your site as well as other sites that are on the WAS system.  In order to start your capture you can click the Capture Icon under the name you gave your site and it will start capturing the site according to your specifications.  An email is sent to you once the site has finished capturing.  Once the capture has finished WAS provides a variety of options for looking at your captured site.  These include an ability to review the entire captured site and navigate around it as if it were live as well as an ability to search the site by keyword while narrowing your search results by file type (such as pdf, images, audio, video, etc.).  Once you have captured your site more than once there are additional options available to compare the results of the captures to see what items have changed or been removed from the site.


In my case, I chose the 36 hour capture option (recommended for first time captures).  Once the capture was done, I signed back on to navigate the harvested site.  The harvested website, as far as I am able to tell, has the highest fidelity to the original site (compared to the other web capturing software I evaluated).   I had no problems running the program or setting up the capture and the documentation was clear and easy to understand.  Considering the ease of use, the quality of the result, and the options for navigating and comparing the finished captures, I am recommending this program as the best of the four programs I have tried.  My only concern for the WAS software is whether there are long term preservation issues because  the captured site resides with WAS rather than our own internal server.


Evaluation Criteria:



  • Installation/Configuration/Supported Platforms: Because WAS is run through a web browser there is no installation necessary.  All you need is access to the internet in order to use the service.  Configuration of each capture is minimal.  All you need is a subscription name and password in order to get started.  20/20



  • Functionality/Reliability: There were no problems in running the capture and on an initial look through the captured site it has a very high fidelity to the original live site. 20/20



  • Usability: Incredibly user friendly.  There are very few parameters to adjust, so it is not overwhelming to a novice.  10/10



  • Scalability: I’m not sure how this would work since I was only capturing one site. However, my guess is that if you were trying to capture a large number of websites it might not scale well since the max capturing time is 36 hours.  It seems very well designed to capture specific hosts rather than single broad captures of numerous host sites.  5/10



  • Documentation: There are a number of helpful video tutorials as well as pdf manuals available. The manual and videos were easy to understand and follow.  10/10



  • Interoperability/Metadata support: There is a description section available to add metadata associated with your captured site.  10/10



  • Flexiblity/Customizability: Unfortunately the thing that makes WAS user friendly is also what makes it less flexible in terms of adjustable parameters.  There are only a few parameters that can be adjusted when setting up the crawl so this program may be frustrating to those who have a lot of programming experience who want to be able to make small adjustments.  To balance this though, there is a lot of flexibility in terms of the output of the program since the site can be searched by keyword, file type, can be compared to previous crawls, or can be navigated as if it were live. 7/10



  • License/Support/Sustainability/Community: WAS is part of the University of California Library system and seems to be very actively used.  There are regular “Introduction to WAS” web conferences held for those interested in learning more about how to use the program.  They provide contact information for service support in case there are issues.  There is also a mailing list, Facebook page, and RSS feed for WAS.  The downside is that it is not an open source program and users outside of the University of California system must pay a yearly fee for use.  While it is likely to be sustainable over the long run since it is part of a large public university system, I do have concerns over the long term preservation of individual sites since they are not available for download (as far as I could tell) from the WAS site to our own internal servers.  8/10


Final Score: 90/100


Bottom Line: If you can afford to subscribe to this service and are okay with not hosting your captured sites in house, this is the program to use.  It is incredibly user friendly, easy to use, and provides output that is of high quality in addition to having features that allow you to compare your captures.

"

Social Media Monitoring Gaining Ground But Has Plenty of Room for Growth

Social Media Monitoring Gaining Ground But Has Plenty of Room for Growth: "

As I prepare this morning to make a presentation on social media monitoring to a group of non-profit executives at the VOLUNTEER Hampton Roads, Hampton Roads Institute for Nonprofit Leadership Conference in Norfolk, VA, eMarketer is sharing an InformationWeek Analytics report about, you guessed it, social media monitoring.


The report shows that there is plenty of room for growth. Of course, we know Andy Beal already gets it (insert blatant Trackur social media monitoring tool plug here) but now it’s up to companies to get it more as well.



With 43% of the respondents having no plan to respond to any online comments it’s obvious that more business people need to understand this process before it is too late. There are many instances of companies and organizations having to scramble to respond to online issues and concerns. While doing something is usually better than doing nothing it is still very risky to have to put together a response on the fly when just some prior planning could prevent some seriously poor performance.


The study also looked at what companies are using to monitor the social space and most are still depending on search engine alerts which, in this space, is quickly becoming the equivalent of using smoke signals to communicate.



Since we are on the subject, be sure to check out the Trackur blog for even more information about the discipline of social media monitoring.


NOTE: One piece of monitoring we would like to do is to see what our readers think of this weekend’s Super Bowl. Who are the Pilgrims rooting for? Let us know in the comments. For me, it’s Go, Pack, Go! What about you?


Pilgrim’s Partners: SponsoredReviews.com – Bloggers earn cash, Advertisers build buzz!




"

Zuckerberg Gets Letter From Congress About Data Privacy Concerns

Zuckerberg Gets Letter From Congress About Data Privacy Concerns: "

If you are a company that depends on your users’ information to make a good portion of your revenue like Facebook does for advertising you likely don’t want letters from politicians about your tactics. It’s like getting a letter from the Principal in school. You know you did something wrong but you are hoping it doesn’t go on your permanent record. Then the letter arrives at the house. Ouch.


Well, Facebook’s Mark Zuckerberg got one of those little notices. It came from U.S. Reps. Edward Markey (D-Mass) and Joe Barton (R-Texas), Co-Chairmen of the House Bi-Partisan Privacy Caucus and it was dated February 2. The concerns come from the plan that Facebook launched in January then pulled off the table to be tweaked for re-release that gave developers access to mobile phone numbers and addresses of Facebook accounts. Representative Markey’s website tells some more.


“Facebook needs to protect the personal information of its users to ensure that Facebook doesn’t become Phonebook,” said Rep. Markey. “That’s why I am requesting responses to these questions to better understand Facebook’s practices regarding possible access to users’ personal information by third parties. This is sensitive data and needs to be protected.”


“Facebook’s popularity has made it a leader in innovation and we hope they will also be a leader in privacy protection,” said Rep. Barton. “The computer – especially with sites like Facebook – is now a virtual front door to your house allowing people access to your personal information. You deserve to look through the peep hole and decide who you are letting in.”


Some of the questions that Facebook is being asked to answer include:


Would any user information in addition to address and mobile phone number be shared with third party application developers under the feature as originally planned, and was any of this information shared prior to Facebook’s announcement that it would suspend implementation of the feature?


What user information will be shared with third party application developers once the feature is re-enabled?


What was Facebook’s process for developing and vetting the feature referenced above before the feature was suspended, and what was the process that led Facebook to decide to suspend the rollout of this feature? What is the process Facebook is currently employing to adjust the feature prior to re-enabling it?


What are the internal policies and procedures for ensuring that new features developed by Facebook comply with Facebook’s own privacy policy, and does the company consider this a material change to its privacy policy?

What consideration was given to risks to children and teenagers posed by enabling third parties access to their home addresses and mobile phone numbers through Facebook when designing the new feature?


What are the opt-in and opt-opt option for this new feature?


Why is Facebook, after previously acknowledging in a letter to Reps. Markey and Barton that sharing a Facebook User ID could raise user concerns, subsequently considering sharing access to even more sensitive personal information such as home addresses and phone numbers to third parties?


It’s the last question in the previous quote that pretty much sums up how Zuckerberg and Facebook approach the world in most cases. You see, back October the company had told these same two representatives that sharing Facebook User ID’s with developers raised privacy concerns. Now, Facebook goes ahead and gets caught with its hand in the privacy cookie jar looking to give away even more sensitive information like mobile numbers and addresses. That’s either chutzpah or just plain disregard for concerns that had been voiced by these reps in the past. If you are Facebook does it make sense to ‘poke the bear’ and open this door again after ticking off the same two men you had issue with in just the past few months?


At any rate, it’s obvious that Facebook will push every envelope it can to get user data in the hands of people that can help Facebook make money. As a capitalist that makes sense. But Facebook’s apparent disregard for any convention of decency is going to come back to bite them in the end. Despite all the accolades and platitudes cast upon Zuckerberg as a visionary etc, etc there is no denying that he is also arrogant and feels to be above the law in many ways. In a word, there are times where he is just unlikeable. This culture and attitude has been created at Facebook as well so this will likely not be the last time the company gets called into the principal’s office.


If you would like to see a full copy of the letter you can do that here.





"

Harry Potter on Electronic Records Management – aiim.org/training | Document Technology News & Info

Good and interesting presentation!

Harry Potter on Electronic Records Management – aiim.org/training | Document Technology News & Info

Thursday, February 3, 2011

AOTUS: Collector in Chief | A National Archives of the Future

"At the National Archives, we are meeting the President’s call to action. Charting the Course is our plan for reinventing the National Archives to meet the demands we face in the digital age".

AOTUS: Collector in Chief | A National Archives of the Future: "At the National Archives, we are meeting the President’s call to action. Charting the Course is our plan for reinventing the National Archives to meet the demands we face in the digital age."

Business English Technology Vocabulary for IT - Web 2.0

This is part 2 and once again an excellent overview of the basics!



English Vocabulary for ESL - IT & Computing: Web 2.0 (Pt. 1)

Good presentation w/the basics!