Friday, 14 June 2013

How to setup ssh environment with a password bashed key pairs

1. Background:

The purpose of this document is to overcome the issue that we face at our AWS environment. Its true that amazon not keeping our private key, when we have create a new instance. But we are not very strict about those aws keys, most of the team inside the company use the key and keep the same in lots of the place, and this can cause of a security hole.

Example:

Let say foo.com is an online company and the qa environment people is not keeping track of where the keys are keeping. They are not very sure about the keys. Let say if any of the qa team person kept the key at document root and later some how that key got into hand of some cracker, then he/she can log-in to the QA environment and that's compromise our environment.

2. What we can do?

Its true that without the key no one can able to log-in. But we don't even want to share those aws keys to everyone.

So, create your own ssh-key pairs, with a pass-phase.

1. ssh-keygen [enter]
2. select your preferred key type [rsa / dsa] or can go for the default one. [ enter ] 
3. in pass-phase enter a password. 

After that you can update the public key of that key pairs to qa server's authorized_keys and share the QA team. Now onwards when ever they want to login to the qa servers, they can use the same key from a blessed host. [ You can create a secure host, from where every-one login to company servers]. And you can protect the blessed hosts in the same way.

3. Question(s)?

How do I perform the automation:

You can do the following for that:

1. Log in to bless host [ with your personal keys ].
2. use "ssh-agent bash" [ I am using bash, you can use any of your shell ]
3. ssh-add [ at this time it will ask for the password, provide the password. ]

Later you can log-in to the qa server from the bless hosts without typing the password again and again.

You might also like to go through "screen" command.

4. Further reading:
 
Please go through further documentation on the following command for more details:

ssh-keygen
ssh-agent
ssh-add
screen


Wednesday, 12 June 2013

Mysql Notes1


1. mysql errcode:24

 mysqldump: Couldn't execute 'show fields from `tablename`': Out of resources when opening file './databasename/tablename#P#p125.MYD' (Errcode: 24) (23)

For above kind of problem, you might need to check the ulimit value of system and mysql user.

- login as mysql user: [ sudo -u mysql bash ]
- ulimit -a [ check for the ulimit values ]
- update the /etc/security/limits.conf

If you find the ulimit is less for the mysql user, you can update the same.

# sudo lsof -p <pid_of_mysql> | wc -l  
The above command will let you know how many file is getting used by the mysql process.

mysql> show variables like 'open%';


2. To delete unwanted mysql database partition:

for i in {n1..n2}; do mysql -uroot --pPassword -e "use database_name; alter table tablename_x drop partition p$i;" ;done

In the above base script, n1 and n2 are starting of unwanted mysql database partition number and end of unwanted mysql databse partition number.


To recheck there should not have unwanted mysql table partition:
show create table table_name_x \G



web_accelerator

I am trying understand some open web accelerator, and as of now going through the data from http://en.wikipedia.org/wiki/Web_accelerator  and found following is the order: if I have to chose a open web accelerator:

1. Varnish: http://en.wikipedia.org/wiki/Varnish_(software)
2. Nginx: http://en.wikipedia.org/wiki/Nginx
3. Squid: http://en.wikipedia.org/wiki/Squid_(software)
4. trafficServer: http://trafficserver.apache.org/  [ ATS: Apache Traffic Server ]

Technologies

NOTE:


www.apache.org have so many new things to learn. Follow the new stuffs over there and you can learn a lots, that is new to the IT.


Few more new technologies:
- Apache Mesos -> http://mesos.apache.org/
  Making it easy to build resource-efficient distributed systems

- Spark -> http://spark-project.org/ :

Spark is an open source cluster computing system that aims to make data analytics fast — both fast to run and fast to write.
Spark is also the engine behind Shark, a fully Apache Hive-compatible data warehousing system that can run 100x faster than Hive.To run programs faster, Spark provides primitives for in-memory cluster computing

http://spark-project.org/examples/

Few top Sites:
https://amplab.cs.berkeley.edu/projects/
http://www.apache.org

Spark, 
Shark, 
GraphX, 
ZooKeeper
https://github.com/amplab/shark/wiki
http://spark-project.org/docs/latest/
https://github.com/amplab/shark/wiki
https://amplab.cs.berkeley.edu/projects/
http://www.scala-lang.org/







•High quality technical documentation, runbooks, diagrams.
•PCI-DSS compliance and implementation.
•Platforms - CentOS/RHEL, Debian, Solaris and FreeBSD.
•Shell scripting, Automation, UNIX/GNU tools.
•Deep understanding of critical networking fundamentals, tools, protocols.
•Virtualization - VMWare Vsphere, ESX/ESXi, Xen.
•Security - Bluecoat, Checkpoint, ASA/PIX, SonicWall.
•Hardware - Dell, HP, Sun, Cisco, IBM, EMC, Netapp, Hitachi
•MTA/SMTP - Postfix/Sendmail, Exchange, Zimbra
•Enterprise Monitoring - Nagios, SolarWinds, Cacti, BMC.
•ITIL - problem/incident management, change management, continual service improvement.

Specialties: Enterprise Systems Administration, Infrastructure Design/Architecture, Automation, HA/DR, Linux, UNIX, Scripting, Bash, Perl, Python, Apache, MySQL, SAN/NAS, Virtualization, LDAP, Virtualization, SMTP, Postfix, VMWare, Xen, Enterprise Server Hardware (HP, Dell, Sun, Cisco, IBM), Enterprise Monitoring (Nagios, BMC Patrol, SolarWinds Orion, Zenoss, Cacti), PCI-DSS, Bluecoat, Netscaler/F5, Active Directory, Exchange, DLP, iptables, MSSQL



*) Technology:

Zabbix
glusterfs
monit
bacula : -> open source network backup.
powerdns

nagios
puppet
Cacti
OpenTSDB
ganglia
Flume
Hadoop
logstash
graphite

Technology I need to learn:
GFS, BigTable, MapReduce, Chubby and large-scale 'cloud computing' clusters.
Languages: Python, Ruby, PHP, Perl, Javascript, Shell
SQL/Database : MySQL, SQLite, Cassandra
Distributed Caches: Redis, Memcache
Distributed Processing: Hadoop + Pig + ZooKeeper + Mahout
Cloud Platforms: Amazon Web Services, Google App Engine, Microsoft Azure
Protocols: XMPP, Jingle, ICE, RTSP, SMTP, POP, IMAP
Installers: NSIS
Code Repository Tools: Git, SVN, CVS
Collaboration: JIRA, Confluence
Build Management: Ant, Maven, MsBuild, Nant, Glu, Hudsan
OS: Linux (Redhat, CentOS), Ubuntu, freebsd
Monitoring: Nagios, Cactii, Ganglia, nagiosgraph, RRDTools, Ntop
CI: Teamcity, Clover, Hudson
Config Management: Puppet
Infrastructure: ServerIron Load Balancers,
File Systems: Ext3, NTFS, XFS, GFS
Mail Servers: Qmail, Postfix, Sendmail
App/Web Servers: Tomcat, Resin, IIS, PHP Accelerator, Jetty, apache http
Mailing List: Exmlm, Mailman, Sympa
Jabber Servers: eJabberd, Tigase, Openfire
VOIP Servers: Asterisk
DNS Servers: BIND, Power DNS,
Proxy servers: Squid, ISA, Perdition, NGinx, Varnish, Perlbal
Few more: tcp/ip, http, load balancers, web servers, memcache,
DB Replication: Slony, MSSQL Replication
FTP Servers: Proftpd, VSftpd
Virtualization: Xen, VmWare
Patch Management: WSUS, Yum, up2date
Apache, memcached, Squid, MySQL, NFS, DHCP, NTP, SSH, DNS, and SNMP
Advanced knowledge of Linux, TCP/IP and web services
A strong background in internet service deployment, provisioning, IP networking, service infrastructure, or software deployments.
Bootstrap, other, awk, sed, tc, cfengine, openNMS, MRTC, OpenVPN, HAProxy.
Security and venerability tools, CISSP, CEH tools.
Storage. 
Dbugging tools, gd, firefox web tools: greasemonkey and other imp web imp tools/plugins.

Queue system: [ Message Broker ]
. Gearman
. RabbitMQ
. ZeroMQ
. Amazon SQS
. WebSphere  MQ [ MQSeries ]
. Advanced Message Queuing Protocol [ AMQP ]
. http://en.wikipedia.org/wiki/Gearman
. Redis
- web analytics softwares:  [ Analog, AWStats,CrawlTrack, Open Web Analytics, Piwik, W3Perl, Webalizer, syslog-ng ]
- awstat [ sudo apt-get install awstat ] at browser: file:///usr/share/doc/awstats/html/index.html
- webalizer [ web log analysis software ]
- nodejs
- mongodb

Development
Languages: Scala, Python, Ruby, Java, C#, VB.net, PHP, VC++, C++, Perl, XUL, Javascript, C, Shell
Web Technologies: HTML 5, CSS, Dojo, jQuery, YUI, Flash, Silverlight
Frameworks & Libraries: Hibernate, Hibernate Shards, Spring, Apache MINA, Project Grizzly, log4j, XAPool, Poolman
RDBMS: Postgres, MySQL, Microsft SQL Server, Firebird, SQLite
NoSql Stores: Redis, Cassandra, Voldemort, Berkeley DB
Distributed Caches: Redis, Memcache
Distributed Queues: Kestrel, RabbitMQ
Distributed Processing: Hadoop + Pig + ZooKeeper + Mahout
Cloud Platforms: Amazon Web Services, Google App Engine, Microsoft Azure
Protocols: XMPP, Jingle, ICE, RTSP, SMTP, POP, IMAP
Scripting: Perl, Python, Ruby
Unit Testing: JUnit, NUnit, MbUnit
Stress Testing: Jmeter, Tsung, Iozone, Iometer, Bonnie, Bonnie++
Functional Testing: Watir, Selenium
Installers: NSIS
Code Repository Tools: Git, SVN, CVS
Collaboration: JIRA, Confluence
Build Management: Ant, Maven, MsBuild, Nant
CI: Teamcity, Clover, Hudson
IDEs: Aptana, Komodo, Eclipse, IntelliJ, Visual Studio, EMacs!
System Administration
OS: Linux (Redhat, CentOS), Windows
Monitoring: Nagios, Cactii, Ganglia
Config Management: Puppet
Infrastructure: ServerIron Load Balancers, Cisco ASA Firewall, FC/iSCSI SANs (Comet, Dell)
Scripting: Bash, Perl, Expect, Python, PHP, VBS, Powershell
File Systems: Ext3, NTFS, XFS, GFS
Other: DRBD, Heartbeat, ldirectord, RIS, LVS
Servers
App/Web Servers: Tomcat, Resin, IIS, PHP Accelerator, Jetty
Mail Servers: Qmail, Postfix, Sendmail
Mailing List: Exmlm, Mailman, Sympa
Antivirus / Antispam: clamd, Razor, Kaspersky server, Pyzor, Policyd, RBL/DNSBL
Jabber Servers: eJabberd, Tigase, Openfire
VOIP Servers: Asterisk
DNS Servers: BIND, Power DNS, DLZ, Microsoft DNS
Proxy servers: Squid, ISA, Perdition, NGinx, Varnish, Perlbal
DB Replication: Slony, MSSQL Replication
FTP Servers: Proftpd, VSftpd
Virtualization: Xen, VmWare
Patch Management: WSUS, Yum, up2date
UI
UI Prototyping: Balsamiq, Axure
Design: Photoshop, Flash, Coreldraw
Web: ECMAscript (actionscript/javascript), RSS, XML, HTML (4.01), XHTML, CSS1.0 & CSS2.1


Few companies are looking for the following skills for there "Site Reliability Engineer" (SRE) :

real-time metrics
deployment
Apache Http
Nginx
mysql
node.js
solr
Ruby
perl
python
java
php
Lucene
Scala
system level thinking ability
SOLR or Lucene
Scala
Performance Tuning
chef
puppet
unix
Scrum
FreeBSD
Ubuntu
Redhat
Hadoop
Riak
ZooKeeper
RabbitMQ
Haproxy
AWS
OpenStack
Cloud Foundry
Heroku
Linux Kernel Internals
Filesystems
High availability
High performance
High security
Salt
SVN, GIT
PostgreSQL
Cassandra
Mysql
tcpdump
device drivers
FreeBSD
dtrace
ktrace
load balancing
LAMP stack
column stores
erlang
haskell
scala
scheme
tcp/ip
http
security
storage
memcache
code-review
Sharing and Guding others
capistrano
mesos
scribe
capacity planing
CDN
memcached
squid
nfs
dhcp
ntp
ssh
dns
snmp
Varnish
Redis
Nagios/Icinga
OpenTSDB

Artifactory Repository

Linux_interview_questions_2

link: http://isitup.wordpress.com/tag/linux-administrator-interview-questions/

Bash:
  1. How do you find out if a shell command succeeded or not?
  2. Write a command line to delete files in /tmp across say 400 machines.
  3. What do $?, $!, $@ stand for?
  4. Write a shell script that replaces the shell command which.
  5. What is your favorite shell and why?
  6. What does . mean in regular expressions?
  7. Write a regular expression that does that matches this
  8. What is the output of ls -l ?
  9. What does the stat command do?
  10. What is difference between ” and ‘  ?
  11. How will you change root password on 10K servers and you have no ssh keys? What will you do if it is 100K servers?  [ANS: http://amitmund.blogspot.in/2013/10/ssh-tips.html ]
Databases:
  1. What are  the differences between innodb vs myisam?
  2. How can you tell if replication is functional?
  3. What does sharding mean to you?
  4. What does explain mean to you?
  5. What are the general parameters you tune for MySQL?
  6. What can you do to shutdown a Oracle database that you cannot sqlplus into?
DNS:
  1. What is the difference between A record and CNAME record?
  2. Why PowerDNS over BIND?
Editors:
  1. Emacs vs VI what’s your choice and why?
  2. Wait, what you use nano?
Configuration Management:
  1. What is puppet? Why do you use it?
  2. What do you need to do  rebuild a puppet master on the fly?
  3. What is the keyword for forcing one-operation after another operation in a puppet manifest?
Mail:
  1. Describe a SMTP transaction with commands.
Networking:
  1. What is the difference between TCP vs UDP?
  2. How  do ping and traceroute work? Do get into details about TTL.
  3. Can you tell me the command line for tcpdump to capture traffic on port 22 on interface eth0?
  4. Can you describe how routing works?
  5. Describe a HTTP transaction in all detail you can.
  6. What is a VLAN? Why do you need VLANS?
  7. What are the port numbers for pop,imap,http,https,dns?
  8. What is a spanning tree?
  9. What are MSS and MTU?
  10. Explain when you would use UDP and when you would use TCP.
  11. What is supernetting?
  12. How many bits in an IPV4 address?
  13. What is the difference between Class address and CIDR?
  14. Describe a DNS transaction in great detail.
NFS:
  1. What is the possible reason that a nfs server would appear to be slow when serving large files but quite fast when serving small files?
  2. What is your experience with nfs?
Operations:
  1. Do hostnames need a sequence number in their naming convention? Why/Why not?
  2. How do you deal with multiple issues that come up?
  3. How do you deal with developers?
  4. How do you handle aggregating logs files?
  5. Describe whats your strategy for disaster recovery?
  6. A client says he/she is expected to receive 5million hits, what hardware do you recommend?
  7. What Linux distribution would recommend for a business?
  8. Do you have any issues with running Windows in production?
  9. Describe your ideal job?
  10. Describe what you would like to do?
  11. What are your thoughts on how a release process should be?
  12. What are your thoughts on monitoring?
  13. Design a high level detail for a LAMP stack and DR solution.
Operating Systems/Distributions
  1. What’s your favorite Linux distro and why?
  2. How do you go about recommending what distro to use for a business?
  3. What is an inode?
  4. What is the difference between a soft link and a hard link?
  5. Can you create a directory/filename by the name -f ? If so how?
  6. What is uptime?
  7. What is load average? What do those numbers mean?
  8. What does it mean when someone says there is high load?
  9. When can you have such a high load of 400 and yet have low CPU usage?
  10. What is a semaphore?
  11. What is is a signal?
  12. Can you name few signals?
  13. What does set group id for a directory mean?
  14. What are the different kinds of files?
Security :
  1. How do you secure a linux system?
  2. What is a sql injection attack?
  3. What is the command to manipulate kernel level firewall?
Misc:
  1. Where do you get your news?
  2. What are you favorite set of tools you like to have always?
  3. Tell us a joke.

External_Good_Links

http://stackoverflow.com/

http://serverfault.com/

http://meta.stackoverflow.com

http://superuser.com

http://programmers.stackexchange.com

http://webapps.stackexchange.com/

http://security.stackexchange.com/

http://dba.stackexchange.com/

http://askubuntu.com/

http://unix.stackexchange.com/

http://stackexchange.com/

http://www.michael-noll.com/

http://www.thegeekstuff.com/

http://graphite.wikidot.com/

http://graphite.wikidot.com/faq

http://graphite.wikidot.com/start 

http://graphite.readthedocs.org 

http://tompurl.com/2011/08/12/installing-graphite-on-ubuntu-10-4-lts/

http://www.smilecouple.org/2011/03/01/fix-out-of-resource-problem-with-mysql 

http://kaivanov.blogspot.in/

http://www.americanscientist.org/issues/pub/2002/3/the-easiest-hard-problem/2

http://www.slideshare.net/sudhirpg/ganglia-monitoring-tool

http://www.ibm.com/developerworks/library/l-ganglia-nagios-1/

http://blogs.hbr.org/cs/2013/02/write_e-mails_that_people_wont.html

http://redis.io/

http://redis.io/topics/data-types#strings

http://redis.io/documentation

http://try.redis.io/

http://redis.io/commands

Hadoop Operations:  [http://www.amazon.com/Hadoop-Operations-Eric-Sammer/dp/1449327052/ref=pd_sim_b_1 ]

http://www.tummy.com/articles/isolating-heavy-load/

http://www.tummy.com/

http://cherry.world.edoors.com/

http://www.infoworld.com/d/application-development/10-programming-languages-could-shake-it-181548

AWS_Amazon_Spot_Instance

AWS: Amazon spot instance:

The purpose of using an Amazon Spot instance is for:
1. To save money (it’s cheaper to run servers on Amazon as Spot instances) 
2. To make recovering a failed instance very fast and easy.

Q1. How this is cheaper?
Few example: 

http://www.youtube.com/embed/WD9N73F3Fao?rel=0&hd=1
http://www.youtube.com/embed/BD1X5ItelOk?rel=0&hd=1

Because, even if we have bid for the higher price, its charges for the current spot price. 

Example: 
1. A linux c1xl server's current bid price is $0.070 [ 7 cents only per hour ]
2. The On demand price is around $0.50 [ 50 cents per hour]

and out bid price is $0.75 [ 75 cents per hour ] which is higher even the on demand price, but how its can cost less? Because its charge on the current spot price [ which is 7 cents now]. Then why we are requesting for that high bid price [ around 75 cents ] even higher then spot price?

Note that, the spot instance goes away if some one bid in higher price, so when we have our bid price is more then the on-demand price, then most likely we have a higher chances that out hosts will not go down, and the spot price stay much lower then on demand price for a longer time. Hence we save an overall money and a higher chances of getting the spot instance for longer time.

For more details, please follow the above youtube links.

A Spot Instance on Amazon is how customers can ‘bid’ for unused capacity on Amazon’s infrastructure. The cost to run a Spot Instance is always fluctuating. As long as that cost is below the maximum ‘bid’ price that we bid, then the server continues to run.  However, when Amazon has system issues, network issues, etc, then the current Spot Instance price will go up, exceeding our maximum bid price. When this happens, the server will be terminated instantly without any warning.  Therefore, Spot Instances are useless for things like a Database Server. And they are only useful for areas of the product that are designed to take advantage of them.

AWS wiki link: http://aws.amazon.com/ec2/spot-instances/ [ Further details ]




Spot Instance Issues
A Spot Instance has to be set up in a special way in order for it to work correctly.

When our maximum ‘bid’ price is exceeded, the server will be terminated. Then, when the current price drops back below our maximum bid price, the server will be re-launched. When it is re-launched, it’s like it was launched for the first time: it gets a new InstanceId, all data is lost, etc. Therefore, any data that must be retained has to be on a separate EBS volume.

Also, the IP addresses that the spot instance uses will also change when it is re-launched. This causes issues for both monitoring and for database access, since the server’s IP address is used for both.

This means that the Spot Instance must have boot-up scripts that:
- mount an /ebs volume where the data is stored,
- all data that we change on a regular basis (Apache Document Root, etc) must be links from the root disk to locations under /ebs
- we must assign an Elastic IP address when the instance boots up, so it gets the same public IP address and Amazon DNS name each time it is re-launched. The Amazon DNS name will not change, but the Private IP Address that it resolves to *does* change.