Monday, July 2, 2012

What is Pruning Links


Pruning Links

Once you have identified if bad links are pointing at your site, you need to start working to address it.
link pruningIf you have 10,000 or more links to your site, this can seem like an insurmountable task – particularly if the people who acquired the bad links are no longer around to ask about what they did.
Here are some things you can do to simplify the cleanup process.

Categorize Your Links

Start by pulling the link data. At a minimum, pull it from Google Webmaster Tools, because that is what Google is reporting.
If you can, it is also great to get link data from Open Site Explorer and Majestic SEO. Integrate all this data into one master list.
Combining your data (and de-duping the results) gives you the largest possible list of links, as individual link data sources can only provide a sampling of the links to your site.
Once you have your master list, it's time to start simplifying the task a bit.

1. Sort Your Links by the Linking URL

Group the links by domain. If a domain links to you from 834 pages, just check 1-3 at most – chances are that if the links are bad on any of the pages, it's bad on all of them.
This by itself is a huge simplification of the task. For example, look at these example link counts:

The blog has 201,048 links from 778 domains. So instead of checking more than 200,000 pages, you only need to check 778 domains. The workload seems a lot lighter already, doesn't it?

2. Separate Links From Blogs

This is hard to do without doing a little programming, but the required programming is easy.
Look for URLs that have "blog" as part of the URL, or load the pages and see if you can find the string "Wordpress", "Moveable Type" and other such blog platforms on the page. This won't give you a complete list of blogs, but it will identify a lot of them for you.
You will need to look at these posts, but you can give the person doing the work simplified guidelines, such as telling them to look for:
  • Low quality posts.
  • In context links with rich anchor text.
  • Multiple links per post.
Also, if 1,100 blogs link to your site, and you look at 100 of them and see significant problems, you know you need to check them all (unfortunately!).
But, if you look at 200 or more and they are all clean, you might start to think you don't need to look at the rest. To be conservative, if you look at half of them, and they are all clean, chances are there was no bad blog campaign underway, and you can skip the rest and focus on looking at other problems.

3. Look for Multiple Links Per Page

This is another hint of a possible problem. Not definitive, of course.
For example, the other a guest post someone did on Forbes had six links to the site of the author, including rich anchor text links in the body of the post. It just looked "off."
link-pruning-multilink-pages
People who buy links tend to be a bit greedy. To them, one rich link anchor text link is good, but several links are great – they want to get their money's worth.
The upside of this greed is it can make it easier for you to recognize potential bad links.
Tracking down pages with multiple links takes little bit of programming, but it isn't too hard.

4. Look for Pages That Link to You With Rich Anchor Text

This is again not a definitive flag, but can focus where you look for trouble.
Consider the inverse rule too – if the only links on the page to you use your URL or business name (assuming that these aren't keyword packed), then chances are that the page in question isn't a problem.
Focus your energy on pages that smell like trouble.

Sunday, July 1, 2012

How to Create Robots.txt and Use of Wildcard and Dollar patters Match


How to Create Robots.txt


The simplest robots.txt file uses two rules:
  • User-agent: the robot the following rule applies to
  • Disallow: the URL you want to block
These two lines are considered a single entry in the file. You can include as many entries as you want. You can include multiple Disallow lines and multiple user-agents in one entry.
Each section in the robots.txt file is separate and does not build upon previous sections. For example:
User-agent: *  Disallow: /folder1/    User-Agent: Googlebot  Disallow: /folder2/  
In this example only the URLs matching /folder2/ would be disallowed for Googlebot.

User-agents and bots

A user-agent is a specific search engine robot. The Web Robots Database lists many common bots. You can set an entry to apply to a specific bot (by listing the name) or you can set it to apply to all bots (by listing an asterisk). An entry that applies to all bots looks like this:
User-agent: *  
Google uses several different bots (user-agents). The bot we use for our web search is Googlebot. Our other bots like Googlebot-Mobile and Googlebot-Image follow rules you set up for Googlebot, but you can set up specific rules for these specific bots as well.

Blocking user-agents

The Disallow line lists the pages you want to block. You can list a specific URL or a pattern. The entry should begin with a forward slash (/).
  • To block the entire site, use a forward slash.
    Disallow: /
  • To block a directory and everything in it, follow the directory name with a forward slash.
    Disallow: /junk-directory/
  • To block a page, list the page.
    Disallow: /private_file.html
  • To remove a specific image from Google Images, add the following:
    User-agent: Googlebot-Image  Disallow: /images/dogs.jpg 
  • To remove all images on your site from Google Images:
    User-agent: Googlebot-Image  Disallow: / 
  • To block files of a specific file type (for example, .gif), use the following:
    User-agent: Googlebot  Disallow: /*.gif$
  • To prevent pages on your site from being crawled, while still displaying AdSense ads on those pages, disallow all bots other than Mediapartners-Google. This keeps the pages from appearing in search results, but allows the Mediapartners-Google robot to analyze the pages to determine the ads to show. The Mediapartners-Google robot doesn't share pages with the other Google user-agents. For example:
    User-agent: *  Disallow: /    User-agent: Mediapartners-Google  Allow: /
Note that directives are case-sensitive. For instance, Disallow: /junk_file.asp would block http://www.example.com/junk_file.asp, but would allow http://www.example.com/Junk_file.asp. Googlebot will ignore white-space (in particular empty lines)and unknown directives in the robots.txt.
Googlebot supports submission of Sitemap files through the robots.txt file.

Pattern matching ( Wildcard and Dollar)

Googlebot (but not all search engines) respects some pattern matching.
  • To match a sequence of characters, use an asterisk (*). For instance, to block access to all subdirectories that begin with private:
    User-agent: Googlebot  Disallow: /private*/
  • To block access to all URLs that include a question mark (?) (more specifically, any URL that begins with your domain name, followed by any string, followed by a question mark, followed by any string):
    User-agent: Googlebot  Disallow: /*?
  • To specify matching the end of a URL, use $. For instance, to block any URLs that end with .xls:
    User-agent: Googlebot   Disallow: /*.xls$
    You can use this pattern matching in combination with the Allow directive. For instance, if a ? indicates a session ID, you may want to exclude all URLs that contain them to ensure Googlebot doesn't crawl duplicate pages. But URLs that end with a ? may be the version of the page that you do want included. For this situation, you can set your robots.txt file as follows:
    User-agent: *  Allow: /*?$  Disallow: /*?
    The Disallow: / *? directive will block any URL that includes a ? (more specifically, it will block any URL that begins with your domain name, followed by any string, followed by a question mark, followed by any string).
    The Allow: /*?$ directive will allow any URL that ends in a ? (more specifically, it will allow any URL that begins with your domain name, followed by a string, followed by a ?, with no characters after the ?).
Save your robots.txt file by downloading the file or copying the contents to a text file and saving as robots.txt. Save the file to the highest-level directory of your site. The robots.txt file must reside in the root of the domain and must be named "robots.txt". A robots.txt file located in a subdirectory isn't valid, as bots only check for this file in the root of the domain. For instance, http://www.example.com/robots.txt is a valid location, but http://www.example.com/mysite/robots.txt is not.

Friday, June 29, 2012

Difference Between HTML Sitemap and XML Sitemap

HTML Sitemap V/S XML Sitemap

An HTML sitemap allows site visitors to easily navigate a website. It is a bulleted outline text version of the site navigation. The anchor text displayed in the outline is linked to the page it references. Site visitors can go to the Sitemap to locate a topic they are unable to find by searching the site or navigating through the site menus.

This Sitemap can also be created in XML format and submitted to search engines so they can crawl the website in a more effective manner. Using the Sitemap, search engines become aware of every page on the site, including any URLs that are not discovered through the normal crawling process used by the engine. Sitemaps are helpful if a site has dynamic content, is new and does not have many links to it, or contains a lot of archived content that is not well-linked.

Which is better: an HTML site map or XML Sitemap? 



Thursday, June 28, 2012

How to Define Canonical in your Page



It sounds like an easy question, doesn’t it? While we hear a lot about duplicate content since the Panda update(s), I’m amazed at how many people are still confused by a much more fundamental question – which URL for any given page is the canonical URL? While the idea of a canonical URL is simple enough, finding it for a large, data-driven site isn’t always so easy. This post will guide you through the process with some common cases that I see every week.

Let’s Play Count the Pages

Before we dive in, let’s cover the biggest misunderstanding that people have about “pages” on their websites. When we think of a page, we often think of a physical file containing code (whether it’s static HTML or script, like a PHP file). To a crawler, a page is any unique URL that it finds. One file could theoretically generate thousands of unique URLs, and every one of those is potentially a “page” in Google’s eyes.
It’s easy to smile and nod and all agree that we understand, but let’s put it to the test. In each of the following scenarios, how many pages does Google see?

(A) “Static” Site

  • www.example.com/
  • www.example.com/store
  • www.example.com/about
  • www.example.com/contact

(B) PHP-based Site

  • www.example.com/index.php
  • www.example.com/store.php
  • www.example.com/about.php
  • www.example.com/contact.php

(C) Single-template Site

  • www.example.com/index.php?page=home
  • www.example.com/index.php?page=store
  • www.example.com/index.php?page=about
  • www.example.com/index.php?page=contact
The answer is (A) 4, (B) 4, and (C) 4. In Google’s eyes, it doesn’t matter whether the pages have extensions (“.php”), the home-page is at the root (“/”) or at index.php, or even if every page is being driven off of one physical template. There are four unique URLs, and that means there are four pages. If Google can crawl them all, they’ll all be indexed (usually).
Let’s dive right into a few examples. Please note: these are just examples. I’m not recommending any of the URL structures in this post as ideal – I’m just trying to help you determine the correct canonical URL for any given situation.

Case 1: Tracking URLs

I’ll start with an easy one. Many sites still use URL parameters to track visitor sessions or links from affiliates. No matter what the parameter is called or which purpose it’s used for, it creates a duplicate for every individual visitor or affiliate. Here are a few examples:
  1. www.example.com/store.php?session=1234
  2. www.example.com/store.php?affiliate=5678
  3. www.example.com/store.php?product=1234&affiliate=5678
In the first two examples, the session and affiliate ID create a copy, in essence, of the main store page. In both of these cases, the proper canonical URL is simply:
  • www.example.com/store.php
The last example is a bit trickier. There, we also have a “product=” parameter that drives the product being displayed. This parameter is essential – it determines the actual content of the page. So, only the “affiliate=” parameter should be stripped out, and the canonical URL is:
  • www.example.com/store.php?product=1234
This is just one of many cases where the canonical URL is NOT the root template or the URL with no parameters. Canonical URLs aren’t always short or pretty – many canonical URLs will have parameters. Again, I’m not arguing that this structure is ideal. I’m just saying that the canonical URL in this case would have to include the “product=” parameter.

Case 2: “Dynamic” URLs

Unfortunately, the word “dynamic” gets thrown around a little too freely – for the purposes of this blog post, I mean any URLs that pass variables to generate unique content. Those variables could look like traditional URL parameters or be embedded as “folders”.
A good example of the kind of URLs I’m talking about are blog post URLs. Take these four:
  1. www.example.com/blog/1234
  2. www.example.com/blog.php?id=1234
  3. www.example.com/blog.php?id=1234&comments=on
  4. www.example.com/blog/20120626
Again, it doesn’t matter whether the URLS have parameters or hide those parameters as virtual folders. All of these URLs use a unique value (either an ID or date) to generate a specific blog post. So what’s the canonical URL here? Obviously, if you canonicalize to “/blog”, you’re going to reduce your entire blog to one page. It’s a bit of a trick question, because the canonical URL could actually be something like this:
  • www.example.com/blog/this-is-a-blog-post
This is why we have such a hard time detecting the proper canonical URLs with automated tools – it really takes a deep knowledge of a site’s architecture and the builder’s intent. Don’t make assumptions based on the URL structure. You have to understand your architecture and crawl paths. If you just start stripping off URL parameters, you could cause an SEO disaster.

Case 3: The Home-page

It might seem strange to put the home page third, but the truth is that the first two cases were probably easier. Part of the problem is that home pages naturally spin out a lot of variations:
  1. www.example.com
  2. www.example.com/
  3. www.example.com/default.html
  4. www.example.com/index.php
  5. www.example.com/index.php?page=about
Add in complications like secure pages (https:), and you can end up multiplying all of these variants. While this is technically true of any page, the problem tends to be more common for the home page, since it’s usually the most linked-to page (both internally and from external sites) by a large margin.
In most cases, the technically correct home-page URL is:
  • http://www.example.com/
…but there are exceptions (such as if you secure your entire site). I don’t see the trailing slash (“/”) causing a ton of problems on home pages these days, since most browsers and crawlers add it automatically, but I think it’s still a best practice to use it.
Another common exception is if your site automatically redirects to another version of the home-page – ASP is notorious about this, and often lands visitors and bots at “index.aspx” or a similar page. While that situation isn’t ideal, you don’t want to cross signals. If the redirect is necessary, then the target of that redirect (i.e. the “index.aspx” URL) should be your canonical URL.
Finally, be very careful about situation #5 – in that case, as I discussed in the first section of this post, the “index.php” code template is actually driving other pages with unique content. Canonicalizing that to the root or to “index.php” could collapse your site to one page in the Google index. That particular scenario is rare these days, but some CMS systems still use it.

Case 4: Product Pages

In some ways, product pages are a lot like the blog-post pages in Case #2, except moreso. You can naturally end up with a lot of variations on an e-commerce site, including:
  1. www.example.com/store.php?id=1234
  2. www.example.com/store/1234
  3. www.example.com/store/this-is-a-product
  4. www.example.com/store.php?id=1234&currency=us
  5. www.example.com/store/1234/red
  6. www.example.com/store/1234/large
If you have a URL like #3, then that’s going to be your canonical URL for the product in most cases (especially #1-#3). If you don’t, then work up the list. In other words, if you have #3, use it; if not, use #2; if not #2, use #1. You have to work with the structure you have.
URLs #4-#6 are a bit trickier. Something like the currency selector in #4 can be very complicated and depends on how those selections are implemented (user selection vs. IP-based geo-location, for example). For Google’s purposes, you would typically want them to use the dominant price for the site’s audience and canonical to the main product URL (#1-#3, depending on the site architecture). Indexing every price variant, unless you have multiple domains, is just going to make your content look thinner.
With #5 and #6, the URL indicates a product variant, let’s say a T-shirt that comes in different colors and sizes. This situation depends a lot on the structure and scope of the content. Technically, your T-shirt in red/large is unique, and yet that page could look “thin” in Google’s eyes. If you have a variant or two for a handful of products, it’s no big deal. If every product has 50 possible combinations, then I think you need to seriously consider canonicalization.

Case 5: Search Pages

Now, the ugliest case of them all – internal search pages. This is a double-edged sword, since Google isn’t a fan of search-within-search (their results landing on your results) in general and these pages tend to spin out of control. Here are some examples:
  1. www.example.com/search.php?topic=1234
  2. www.example.com/search/this-is-a-topic
  3. www.example.com/topic
  4. www.example.com/search.php?topic=1234&page=2
  5. www.example.com/search.php?topic=1234&page=2&sort=desc
  6. www.example.com/search.php?topic=1234&page=2&filter=price
The list, unfortunately, could go on and on. While it’s natural to think that the canonical version should be #1-#3 (depending on your URL structure, just like in Case #4), the trouble is pagination. Pages 2 and beyond of your topic search may appear thin, in some cases, but they return unique results and aren’t technically duplicates. Google’s solutions have changed over time, and their advice can be frustrating, but they currently say to use the rel=prev/next tags. Put simply, these tags tell Google that the pages are part of a series.
In cases like #5-#6, Google recommends you use rel=prev/next for the pagination but then a canonical tag for the “&page=2” version (to collapse the sorts and filters). Implementing this properly is very complicated and well beyond the scope of this post, but the main point is that you should not canonicalize all of your search pages to page 1. Adam Audette has an excellent post on pagination that demonstrates just how tricky this topic is.

Know Your Crawl Paths

Finally, an important reminder – the most important canonical signal is usually your internal links. If you use the canonical tag to point to one version of a URL, but then every internal link uses a different version, you’re sending a mixed signal and using the tag as a band-aid. The canonical URL should actually be canonical in practice – use it consistently. If you’re an outside SEO coming into a new site, make sure you understand the crawl paths first, before you go and add a bunch of tags. Don’t create a mess on top of a mess.

Tuesday, June 26, 2012

Google Panda Algorithm Update 3.8 on 25 June 2012


Google Panda Update 3.8


Google has announced an update with the new Panda algorithm was pushed to a recent day. According to the message from Google in Twitter update "will be noticeable affects only ~ 1% of queries worldwide."
Google Panda Update 3.8So over the weekend negotiations an update, but search giant announced the rollout was officially introduced on 25 June. And there are no updates or changes to the algorithm as it is just a basic data refresh.

The last update was pushed out Panda on 8 June and earlier 26 april. Short, Google does Panda Penguin algorithms and updates approximately every month. But the recent update of Panda took about 2 weeks ago. No doubt, Google'd a few new ideas and wanted to push a new refresh.

Anyway, the "fairness warriors" against plagiarism happy to hear news. The new Panda update is another issue for webmasters force them to create unique and quality content. But webmasters are afraid that the new algorithm will have negative impact on their websites. Perhaps some would respect Google as the search giant more used things like Panda to his own pages. What do you think of it?

Monday, June 25, 2012

How to Beat Google Penguin Update

Google’s Penguin update was meant to punish those people who do lazy SEO work in an attempt to climb the SERP by tricking the big G. There are a few things that come into play here…

Google Penguin update


What is Perfect keyword density of a page? - Matt Cutts