Senin

Index search


Search index
is a database search engine, in which a structured fine saves and
stores information about Web documents collected by search engine
crawlers.
The process of adding documents to the database search system, organizing and storing them is called indexing. Each search engine indexing performed by their algorithms.

What is included in the index

The content of the index is structured data consisting of key phrases,
text and multimedia elements, links, so that shortened the time it takes
to search the database and find the documents in accordance with the
request.
This database is constantly updated using crawlers that continuously scan the web for new pages, content and resources.

Adding a site to the search engines
In the process of positioning the indexing plays an important role. The faster service is indexed by a bot indexing, the sooner you can expect that it will appear in search results. To speed up the indexing, you must notify the search engine of the existence of a new page in the network. For this purpose, all search engines exist function (adurl), where by filling out a special form pages are added to the index. The positive impact can also create a profile page on social networking and dissemination of information in their website.
Management is indexed pages using robots.txt.

HTML


(HTML stands for Hyper Text Markup Language) - hypertext markup language on the Internet. With the ability to interpret HTML browsers can define the appearance of the document to be displayed in them. This language is widely used in the world and most commonly used to create web pages.


Documents written using HTML are processed by web browsers that
understand the language of signs and transform them into easy to listen
format.
Typically, these files have the extension htm. or html. For editing HTML documents, you can use any text editor such as Notepad. There are also designed specifically for this purpose programs, such as Adobe Dreamweaver.


History of HTML


HTML was created by Tim Berners-Lee in the 90's of last century. It was originally only used as a tool to create and share documents by users.
The invention was revolutionary: the user, using the reference, your
computer could see the documents that were found in any other location
in the world.
The main task of the first version of HTML was correct text playback on different devices without any structural distortion.
Since then, HTML has been repeatedly modified and significantly expanded the ability to view documents. This language has several versions:

The first versions of HTML in the early 90's do not have a detailed
specification, because at that time there was no single official
language standard.

The purpose of the first version of HTML was the only word processing
and the use of the most popular styles of formatting such as bold test,
italic, etc.

HTML 2.0 - Added ability to handle forms.

HTML 3.2 - the ability to create tables, display mathematical formulas,
graphs, text wrap effect and the passage of the elements.

HTML 4.0 - Some HTML elements have been removed and in its place proposed to use CSS tables. Added support for scripting and frames.
HTML 4.01 - Modified version 4.0.
HTML 5 - year 2010 - until today. The fifth version is currently under development, work on it should be completed in 2014.

Format HTML document

All HTML documents are created using tags - special tags, most often occurring in tandem - the opening and closing tag. Among them is the attribute and content or only content.
Top element (opening tag) is defined as an attribute should be
formulated within, and finally - where the formatting should be
terminated.


Sample text bold by tag: b
<b> bold text </ b>
already be formatted as follows:
bold text

Standard HTML document contains a mandatory set of tags and has the form:


<! DOCTYPE HTML PUBLIC "- / / W3C / / DTD HTML 4.01 / / EN"
" http://www.w3.org/TR/html4/strict.dtd ">
<HTML>
<HEAD>
<TITLE> Document Title </ TITLE>
</ HEAD>
<BODY>
The text of the document
</ BODY>
</ HTML>

where:
XHTML 1.0 Transitional / / EN "- reveals the nature of the document.
<html> - Beginning of the document;
<head> - use the page header, it posted the information for browsers and search engine spiders;
<title> </ title> - the title of the document is indexed by the search engines;
</ Head> - the end of the page header;
<body> </ body> - the beginning and end of the document, which shall be indicated specific content of the page;
</ Html> - end of html document.

Google


Google - the most popular Internet search engine, owned by Google Inc., an American corporation.

Google's name is derived from the misspelled (specifically or
accidentally) by one of the makers of the English word googol, which
means ten to the hundredth power.
From the name Google has developed a verb to google - guglować, guglić, guglać (all forms are correct), meaning the search for something on the Web using the Google search engine.
Since Google is the most popular search engine in the world, a priority
in the process of positioning is to gain the trust of Google
Positioning page.


The history of the

Google Search was created in March 1996 by Stanford University
students, Larry Page and Sergey Brin, in the work of the Digital Library
Project.

September 15, 1997, he officially domain google.com was registered, and
the following year was created corporation Google, Inc..
At the end of 1998 Google's crawler about 60 million crawled web pages. Google is constantly developing and improving. Search algorithms are updated on average 300 times a year, and the amount of features and services are rising.

Indexing pages
In the process of indexing (scan pages in order to introduce them to the database) Google uses special crawlers:


  • Googlebot - the main robot that scans the content of the page in order to create a search engine's index. - Googlebot-Image - scans the page content to index images.

  • Mediapartners-Google - scans the content of your pages for the presence of AdSense ads.

  • AdsBot-Google - checks for quality content pages that AdWords ads are placed.

  • Googlebot-Mobile - indexed web portals for mobile devices.

  • Google Search Appliance - is responsible for the indexation of the Google Search Appliance devices.



There are two primary theories of indexing:


  1. "Sandbox effect".
    The essence of this phenomenon is that Google places specific pages
    (those that have new domain names, the young sites) in the so-called
    sandbox (waiting area), where the parties are located until the system
    deems them ready to appear in the search results.

  2. The opposite theory, which says that the new party is assigned a high PageRank and high positions in search results. Typically, this privilege working soon - until the real parameters of the test.




In order to avoid placing the works in the Google sandbox, avoid too
rapid acquisition of external links for the "new service" and gradually
acquiring links from trusted sites, so that the process looks natural to
the search engines.


Create a page ranking

One of the most important factors that influence the order of the pages
in the Google search results (as determined on the basis of their
credibility, presumably, stability, etc.) is the PageRank indicator or
simply LP (named after the creator, Larry Page `s - co-founder of
Google) - the number of specifying the value of the by Google.
The PR entire site adopted to recognize LP's main page. PR value can be between 1 and 10 High PR is one of the guarantors of the emergence of the highly in search results. The term PR side affects mainly the quantity and quality of links leading to it. Google did not disclose the exact algorithm for calculating PageRank, and hides data on the frequency of updates. As of today, is one of a number of parameters (it is believed that there are over 200), affecting the position of the hand.
Important is the degree of relevance of the page, the inner contents
and code optimization, quality and quantity of links, and other factors.

Google provides its users the ability to use a variety of services. Among other things, these are:


  • Google Search (search) - the main Google.
    The search is based on the content of the pages also scans files in
    PDF, RTF, Flash SWF, PostScript, Microsoft Word, Microsoft Excel,
    Microsoft Power Point and others.
    There is also a voice search function.

  • Gmail - mailbox.

  • Blogger - service providing the possibility of running a blog.

  • AdSense and AdWords - programs adequately develop and deploy contextual advertising on Google and its network.

  • Google Maps - service with a geographical map of the world.

  • Google Music - gives you the ability to create audioteki on Google's servers.

  • Google Books - Search books (over 10 million. Titles of the largest libraries in the U.S. and elsewhere).

  • Google Checkout - electronic payment system.

  • Google Finance - find financial information for major international corporations.

  • Google Translate - online encyclopedia.

  • Google Scholar - search of scientific publications in a variety of formats, with a wide spectrum of disciplines.

  • Google Video - Search and services of hosted video files.

  • Google News - news service.


FTP

FTP (File Transfer Protocol) - one of the protocols used to exchange data over a network. It is used for loading files, web documents with personal computers (Client), the servers that provide web hosting.

Key phrase

Key phrase
(keyword queries, keys, key phrase) - a word or phrase that is
positioned on the site in search engines, which are also part of the
content Positioning page.
Search results show the most relevant to the user's query websites, and the same search is based on the keywords you entered. Therefore, when creating text for the Positioning is very important to place the optimal number of keywords.
It is recommended, however, that the amount does not exceed the
standards of the search engines and the same text remains readable and
understandable to the user.


The density of the keywords in the text


The density of keywords (called Keyword density) is the ratio of the
amount contained in the text keyword to the number of words.
The density is expressed as a percentage. The optimal keyword density is 3-5%.
For example, the text promoted a preset list of specific keywords and
set a limit of 300 words, the calculation of the key phrases that should
be evenly distributed throughout the text, as follows: if the
recommended density is 3-5%, then we get 300 * 0 , 03 = 9
It follows that in this case you can use keywords nine times. Note, however, that the calculation of densities should not be taken into account stop-words (called stopwords).
It is also worth mentioning that, depending on the frequency of query
and its competitiveness is not only directly make key phrase (exact
match), but the use and form of expression, and extensive descriptions
(use additional words) to zagęczszenie not too high, which could be
treated as spamming by words.


Keywords tag

A web page containing text should have headers. It organizes its logical structure and thus simplifies the collection of text. The header tags H1-H6 necessarily should include the keywords.
The H1 headings should be used immediately, while lower levels in the
headlines - should resort to forms of expression and more rozbudowanch
entries.

Similarly, the title tag should be the location directly key phrases
and word combinations containing them and the surrounding text.
It is necessary to separate a small amount of keywords in the text of logical tags: bold (strong) or italic (em). Mandatory to fill in the fields description, and keywords, where you can enter occurring on the keyword phrases.

Keyword Selection

Among the selected key phrases may be asking for different popularity - HR, PA and NP (high, medium, low popular). For promotion on search engines to be effective you must first properly assess the competitiveness of selected key phrases. Such an assessment will choose the keywords that will be the easiest to promote and bring the highest return. Promoting a page using wysokokonkurencyjnych, common words is difficult, unpopular and may not produce the desired result. The main link key words are on average common queries.

DMOZ


DMOZ, also known as COP (called Open Directory Project) - acclaimed, multi-lingual "white" web directory which undergo a rigorous manual moderation. Catalog edited by a group of enthusiasts, working for free, so anyone willing can contribute to its development. The catalog has been created by U.S. developers in 1998. In the same year, Netscape corporation bought it, which was then acquired by AOL, which owns the directory today. The catalog has a total 5 million pages, and Polish-language section is about 73,000 pages.

Effect on positioning

DMOZ plays a significant role in the process of optimizing your pages,
as it represents a huge database of hand selected high quality services.
Base directory is constantly updated based on Google resources.

Advantages DMOZ directory for positioner:


  • the presence of the catalog shows it as really valuable resource;

  • link from the directory increases the level of confidence of the search engines;

  • positive impact on the growth parameters, such as PR, increase footfall and value of outbound links.



For webmasters - owner of the site - a resource placement in the directory is also important from a financial point of view. Links to pages from the catalog are priced higher than the sides of the services that are not listed in DMOZ.

Content


Website Content (Content) - a general term characterizing any information that can be found at the website.
The literal meaning of the content of the texts can be called, audio
and video files, images, graphics, animations and other information
placed on the site, or anything that you can hand on it to see, hear and
read.

Characteristics page content

The most important feature of providing a quality website is its uniqueness. On this basis, there are two types of content distributed on the Internet:


  • unique content that has never before been published on the Internet (in the search engine index no information about them);

  • non-unique content that have already been deployed in the Internet.



Unique content should have a value of information.
Distribution of unrelated words logically does not meet with a positive
response from both network users and crawlers that have long learned to
recognize similar tricks (such as incompetent multiplication products)
and absolutely banish these pages.

Taking into consideration the manner of its receipt of the content may be:


  • author - deployed on by the owner;

  • created by users (comments, photos, videos distributed on the visitors).



Depending on the contents of distributed content, stands out:



  • thematically related content (for example, an article about the content
    and code optimization for search engines on the web page positioning);

  • content unrelated topic (for example, a film about making the web portal a car).



Moreover, the content can be divided into:


  • solid;

  • refillable (for example, the content of blogs, news sites).


Ways to create content

Taking into account the method of creating content can be conventionally separated into three groups:


  • approved by retrieval systems.
    These are text, graphics, video, etc. created independently by the
    owner, bought the so-called content or exchanges created under the
    relevant orders by experts in the field (most copywriters).
    It is the most effective, but also the most expensive way to create content;

  • tolerated by the search engines.
    The website owner hired a specialist or change existing external
    content to the extent that crawlers found them unique, but the starting
    point remained the same (eg, text editing, change the images in computer
    graphics programs, etc.).

    It is a method not consuming so much time and money as the previous
    one, but the ratio of users to such content is more critical;

  • not accepted by the search engines. This content obtained by copy-paste (copy and paste), or plagiarism other people's texts or images. It is a most undesirable way of creating content.



The importance of the content of the positioning of the

Content is extremely important in the process of positioning the hand
and is one of the key factors affecting the growth of the search engine
page rankings.
For positioning to be effective, content must meet certain requirements:


  • Uniqueness.
    Non_unique content stands in the way of achieving high results, reduces
    confidence in the side, and quite often it is not even indexed by
    search engine spiders.

  • Optimized for search engines (this applies particularly to texts).
    Dedicated to the text keywords and phrases give robots the possibility
    to najdokładaniejszego determine what pages of the site and, as a
    result, views them as best suited to the search results based on
    specific queries.

  • Informacyjność and timeliness. Interesting articles help to attract high-quality audience and keeping her attention fixed.

  • Ease of receipt. The individual elements of the content should be arranged on the page in a logical and understandable.

  • Trust (called trust) (not to be confused with TrustRankiem).
    One of the main contents of the cell used in the positioning of the
    excitation at the user directly to the trust through
    confidence-inspiring content.

    Be sure not to put pressure on the visitors to your website in order to
    obtain the desired result, or do not use an open advertisement.

    You should not feel that he is compelled to do anything, simply direct
    their attention to specific goods (services, events, etc., depending on
    the subject of the page).